AI evaluations that catch real failures: a practical guide to datasets, graders, and regression tests
Turn vague expectations into repeatable tests for AI answers, RAG systems, and agents without trusting one aggregate score.
Reviewed September 30, 2026. The evaluation method is provider-neutral. Product names and hosted evaluation platforms can change; preserve your test cases and grading logic in portable formats.

“The new model feels better” is not an evaluation. Generative systems vary from run to run, and a polished answer can hide a missing fact, a wrong citation, an unsafe tool call, or a process that costs ten times more. Evals turn the desired behavior into repeatable tests.
An eval is not only a public benchmark. For an application, the most useful eval is usually a private set of realistic tasks, expected properties, and graders tied to the product's actual risks.
Start with a decision, not a metric
Every evaluation should support a decision: deploy this prompt, choose between two models, accept a retrieval change, or block a tool from production. Without that decision, a metric can improve while the product does not.
| Decision | Evaluation unit | Useful measures |
|---|---|---|
| Change a support prompt | Realistic support request plus policy context | Correct resolution, policy compliance, escalation quality. |
| Replace an embedding model | Question with relevant and irrelevant passages | Recall at k, ranking quality, downstream answer support. |
| Allow an agent to edit issues | Repository task with tool access | Final repository state, correct tool parameters, unauthorized actions. |
| Reduce cost | The same fixed task set | Success rate, median latency, tokens, and cost per successful task. |
Build the dataset from real work
Begin with 30–50 cases collected from actual questions, traces, bugs, and edge cases. This is large enough to reveal patterns and small enough to inspect manually. Add synthetic cases later to broaden coverage, but do not let generated examples replace observed failures.
Each case should contain:
- a stable identifier and category;
- the user input and only the context available at that point;
- the reference facts, allowed evidence, or expected end state;
- a grading rubric with pass, partial, and fail conditions;
- metadata such as language, difficulty, risk, and source date.
Freeze time-sensitive evidence. If a case asks for the current policy, store the policy version used by the test. Otherwise tomorrow's correct answer may differ from today's reference.
Include ordinary, difficult, and adversarial cases
- Common successful requests
- Ambiguous requests
- Missing information
- Known product failures
- Conflicting evidence
- Unanswerable questions
- Tool errors and timeouts
- Prompt injection and unsafe requests
A test set containing only clean happy paths rewards demos. A set containing only exotic attacks overstates rare behavior. Keep category-level results so the overall average cannot hide a collapse in one important group.
Three grader families
Anthropic's agent-evaluation guide distinguishes code-based, model-based, and human graders. A robust suite usually combines them.
| Grader | Best use | Strength | Main risk |
|---|---|---|---|
| Code | Exact values, schemas, tests, final database or file state. | Fast, cheap, reproducible. | Brittle when many outputs are valid. |
| Model | Semantic support, tone, completeness, rubric-based quality. | Scales nuanced judgment. | Bias, inconsistency, and preference for style over truth. |
| Human | High-impact cases, rubric calibration, disputed outputs. | Contextual judgment and accountability. | Slow, costly, and subject to reviewer disagreement. |
Use code whenever the world provides a check. If an agent is asked to fix a test, run the test. If it must create a record, inspect the record. A model grader should not replace a deterministic outcome check merely because it is easier to call.
Design a model grader like a measurement instrument
Give the grader a narrow rubric, the relevant evidence, and an output schema. Ask it to score one dimension at a time. “Is this answer good?” mixes factuality, relevance, style, and safety into a number that is hard to debug.
{
"criterion": "evidence_support",
"score": 0,
"allowed_scores": [0, 1, 2],
"rubric": {
"0": "major claims lack support or contradict the sources",
"1": "core answer is supported but one material claim is weak",
"2": "every material claim follows from the supplied sources"
},
"evidence": ["document A", "document B"],
"answer": "candidate output"
}
Calibrate the grader against a human-labeled subset. Measure agreement by category, inspect disagreements, and revise the rubric before scaling. Periodically repeat the calibration because model and prompt changes can alter the grader itself.
Evaluate RAG as a pipeline
A RAG answer can fail even when its prose looks reasonable. Separate retrieval from generation:
- Retrieval: did the relevant passage appear in the candidate set?
- Ranking: was useful evidence placed where the context builder could select it?
- Grounding: do material answer claims follow from retrieved evidence?
- Abstention: does the system say when the collection cannot answer?
- Citation accuracy: does each citation support the nearby claim rather than only discuss the same topic?
If retrieval recall is low, changing the answer prompt cannot recover evidence the model never received. Layered measures point to the component that needs work.
Evaluate agents by outcome and trace
Two agents may take different valid paths. Do not require an identical sequence unless the sequence itself is a safety or policy requirement. Grade the final state, then inspect trace properties that matter.
| Layer | Question | Example check |
|---|---|---|
| Outcome | Was the task completed correctly? | Tests pass; ticket fields match the request. |
| Tool use | Were tools and arguments appropriate? | No write tool used for a read-only request. |
| Authority | Did the agent stay inside permission? | No message sent without required approval. |
| Efficiency | Did it avoid loops and needless calls? | Maximum calls and time; repeated-call detector. |
| Recovery | Did it handle tool failure honestly? | Error is surfaced; no fabricated success. |
Run repeated trials when variability matters
One pass can hide nondeterminism. For critical or highly variable cases, run several trials and report a pass rate rather than one result. Keep temperature, model version, tool environment, and prompts recorded so a regression can be reproduced.
Separate the development set from a held-out set. If every prompt change is tuned against the same cases, the application can overfit its eval just like a model overfits training data.
Read the score with uncertainty
A rise from 82% to 84% may reflect only two changed cases in a 100-case set. Inspect which cases changed, whether they are independent, and whether improvements came with regressions elsewhere. Report sample count and category breakdown beside the headline number.
Always inspect failures. The purpose of an eval is not to manufacture a high score. It is to discover which failures remain, how severe they are, and which change caused them.
A minimal regression loop
Define success→Measure
Run fixed cases→Inspect
Classify failures→Improve
Change one layer
- Version the dataset, rubric, prompt, model, and application code.
- Run a baseline before making the change.
- Change one major variable when practical.
- Compare overall and category results, latency, and cost.
- Review new failures and a sample of apparent passes.
- Require a threshold appropriate to the risk before deployment.
NIST's AI risk guidance places evaluation inside broader risk management. That is the right frame: evals measure behavior under defined conditions; they do not prove that a system is safe in every environment.