---
title: "Braintrust | Skills We Assign For | Azendo"
description: "Braintrust for AI evaluation — experiment comparison, scoring, plus the Azendo roles assigned for it."
url: "https://azendo.co/skills/braintrust/"
---

[Skills](https://azendo.co/skills/) AI and machine learning 

# Braintrust.

Braintrust is an evaluation platform for AI applications, providing dataset management, scoring functions, experiment comparison and logging so prompt and model changes can be assessed systematically.

## Where Braintrust fits on a long engagement.

Side-by-side experiment comparison is what the tooling adds over a script. Running several prompt or model variants against the same dataset and seeing where each improved or regressed, case by case, turns a single aggregate score into something actionable.

Aggregate scores hide the important detail. A change that raises the average while breaking a category that matters commercially is a regression dressed as an improvement, and only per-case comparison surfaces that.

## What an assigned team does with Braintrust.

Scoring functions have to encode what actually matters. Exact-match scoring on a task where several phrasings are correct measures the wrong thing, and teams optimise toward whatever the score rewards.

Designing scores that reflect real quality is the substantive work, assigned as explicit scope under [managed ai services](https://azendo.co/services/ai-engineering/).

## What we use Braintrust for.

* Variants compared case by case Where each change helped and hurt, rather than one aggregate number.
* Scores that reflect real quality Criteria designed for the task, because teams optimise toward whatever is measured.
* Regressions visible before release Category-level comparison catching a change that improves the average and breaks what matters.

## How Braintrust capacity is assigned.

Evaluation tooling is assigned inside AI capacity, with scoring design treated as the deliverable rather than the platform setup.

## Roles we assign Braintrust for

* [AI Engineer AI engineering](https://azendo.co/services/ai-engineering/ai-engineer/)

## Service lines it sits in

* [AI engineering](https://azendo.co/services/ai-engineering/)

Capacity is agreed as a committed monthly capacity across a discipline, not per skill.

## Related in ai and machine learning

* [PyTorch — skill we assign for](https://azendo.co/skills/pytorch/)
* [TensorFlow — skill we assign for](https://azendo.co/skills/tensorflow/)
* [scikit-learn — skill we assign for](https://azendo.co/skills/scikit-learn/)
* [Pandas — skill we assign for](https://azendo.co/skills/pandas/)
* [NumPy — skill we assign for](https://azendo.co/skills/numpy/)
* [embeddings — skill we assign for](https://azendo.co/skills/embeddings/)
* [function calling — skill we assign for](https://azendo.co/skills/function-calling/)
* [vLLM — skill we assign for](https://azendo.co/skills/vllm/)

## Tell us what your roadmap needs Braintrust for.

A service delivery manager replies with the disciplines we would assign, the monthly capacity and what the first month looks like.

[All skills we assign for](https://azendo.co/skills/)
