Evaluate

Measure against quality baselines

Build datasets, define scorers, run experiments, and gate deployments. Measure quality and improve iteratively.

Scores and metrics
GPT-5 mini
Base
GPT-5 nano tuned
Comparison
Comparison grade
Regression
Comparison has better accuracy, factuality, helpfulness, instructionFollowing, tone, safety, completeness, citationQuality, latency, cost, errors, and load.
Accuracy
86%avg
92%-6%
3
Factuality
88%avg
93%-5%
2
Helpfulness
82%avg
89%-7%
2
Instruction following
80%avg
90%-10%
3
Tone
87%avg
90%-3%
11
Safety
94%avg
96%-2%
1
Completeness
79%avg
88%-9%
2
Citation quality
73%avg
85%-12%
2
Duration
2.4savg
1.8s-0.6s
2
Total cost
$0.031avg
$0.024-$0.007
2
Total tokens
1,840avg
1,420-420
2
Prompt tokens
820avg
690-130
1
LLM calls
1.5avg
1.25-0.25
1
Errors
0.5avg
0-0.5
1
The eval lifecycle

From “I think this is better” to “I know it is”

01

Build datasets

Golden sets from production logs, user feedback, or manual curation.

Create a dataset
02

Define scorers

Automated, LLM-as-judge, and human review in one place.

Write scorers
03

Run experiments

Compare prompts, models, strategies. Side-by-side diffs.

Run an experiment
04

Gate deployments

CI/CD quality gates block bad changes.

Set up CI gates
Measure quality

Three ways to score. Use them together.

Combine deterministic checks, LLM judges, and human review in one platform. The same scorers run offline in experiments and online in production.

95%Accuracy
62%Helpfulness
Instruction following

Automated scorers

Deterministic checks for format, structure, accuracy.

Build code scorers
Name
Helpfulness
Model
GPT 5 Mini
Score
62%
Rationale

The response directly answers the order status question with clear tracking details and a friendly follow-up offer.

LLM-as-judge

Score tone, helpfulness, reasoning quality.

Configure LLM judges
Response polish

Choose an option to rate the polish level of the response

Human review

SMEs evaluate and build labeled datasets.

Set up review

Compare everything, side by side

Prompts, models, strategies, quantified across every test case.

Experiments

Rich diffs

Run the same dataset through two configs and see what changed.

Prompt v3 compared to Prompt v4
Diff
Quality over time

Track regressions across releases

Every experiment is a data point. See the trend before you ship.

Tone87%
Helpfulness87%
Factuality88%
The complete eval toolkit

Everything you need to measure quality

The feedback loop

Production data improves evals. Better evals improve production.

From logs to labels to eval results to production. Close the loop between what ships and what you measure.

Customer stories

From ad hoc testing to systematic evals

Sarav Bhatia, Sr. Dir. of Engineering

Braintrust is the core of our evaluation framework process.

Josh Clemm, VP of Engineering

We can run hundreds to thousands of experiments with Braintrust.

Mohsen Sardari, VP Engineering

Braintrust helps us ship AI agents customers actually trust.

Start building

Free to start. No credit card required.