Stop Iterating on Vibes: Build an Eval Harness for Your AI Agent
Model, prompt, or tool changes can silently break your agent while you're busy celebrating the fix. This guide shows how to build an eval harness from scratch: a 20-case golden set from real traffic, deterministic code scorers plus calibrated LLM judges, a three-layer scoring split, an anti-self-deception checklist, and a CI gate that actually blocks merges — closing the loop with production sampling and shadow runs.

1. Why "Feels Better" Is Not a Measurement
Every agent builder knows this moment: last night you shortened a wordy sentence in the system prompt, swapped model A for model B, and added one line to a tool's parameter description. Then you tossed three or four questions at the chat box, the answers all looked fine — "yeah, feels better." You hit merge and shipped. Three days later, users start complaining: "It used to file my invoices into the right month. Now it just dumps everything into the current month."
That is iteration without an eval set: flying blind. An agent's capability is a combination of three variables — model, prompt, tools. Change any one of them and capability can quietly collapse along some other dimension. And human intuition about conversation quality is a terrible measuring instrument: you remember the examples that made you go "wow" and automatically forget the old capabilities that regressed. Psychology calls it confirmation bias. Engineering calls it regression — except traditional regression has CI standing guard, and you have nothing.
Fixing A while breaking B is the most common crash pattern in agent iteration. I've seen three recurring scripts:
- Script one: local prompt optimization. You add "always apologize first on refund issues" to your support agent. Refund scenarios improve — but the apology template pushes key parameters out of the context window in multi-turn conversations, and order lookups start dropping order numbers.
- Script two: model upgrade. You move from one strong model to a cheaper or faster one. Eighty percent of tasks look about the same, but the new model is noticeably worse on some dimension you never thought about — say, strict adherence to JSON output format — and tool calls start failing intermittently.
- Script three: tool sprawl. After adding the twelfth tool, tool-selection accuracy slides from 95% to 82%. More tools, more wrong picks. You hand-tested the new tool's scenarios twice and called it good. Nobody ever re-checked call quality on the old tools.
All three share one trait: they happen where you weren't looking. An eval harness exists for exactly one purpose: to turn "I feel like it got better" into a repeatable exam paper. Every change runs the paper first; you read the total and the per-question wins and losses before deciding to merge.
One opinion up front, so we don't argue about it later: evals are not for proving how good your agent is. They are for cheaply discovering where it got worse. The former is marketing material; the latter is an engineering asset. Everything in this method is designed for the latter.
2. Three Layers of Evaluation: One Score Hides Problems
The first trap beginners walk into is giving the agent a single score. "This iteration scored 8.7, last time 8.4 — progress." That number tells you almost nothing: is the 0.3 gain real improvement or judge-model noise? Which capability swallowed the lost points?
The practical decomposition splits agent behavior into three layers, top to bottom. A failure at each layer means something completely different:
Agent 评估的三层拆解
Layer 1: Task success rate (end to end)
One question: did the user's problem get solved? This is the only metric directly tied to user experience. "File this invoice into the correct month and confirm back to me" — the agent runs the whole flow, the invoice lands in the right month, the confirmation comes back. Success. This layer is a black box: however many detours it took internally, a correct outcome counts.
The pass criteria for task success must be operationalized: "solved" can't rest on a judge's vibes. Define it as checkable end states — the file written to the right path, the database row updated, the reply containing the key facts. Which naturally leads to layer two: when a task fails, you need to know where it broke.
Layer 2: Tool-call correctness
Most agent failures live in tool calls: wrong tool chosen, wrong parameters, wrong order, a necessary call skipped. This layer is the most deterministic and the most worth investing in first, because it can be asserted with code — no LLM judge required.
Split tool calls into three sub-metrics:
- Tool selection accuracy: did it call A when it should call A (and not call B when it shouldn't);
- Parameter correctness: right tool, right arguments (types, required fields, enums, date formats — where most failures hide);
- Call sequencing: does it look things up before acting on them? Does it rush to act before the query results arrive?
Layer 3: Output quality
The tools all fired correctly, the task "completed" — but the reply itself can still be wrong: citing documents that don't exist (hallucinated citations), answering a different question than asked, or a tone that annoys the user. This layer is the most subjective, and it's the home turf of LLM-as-judge. Evaluation usually converges on three dimensions: faithfulness (is everything stated grounded?), relevance (did it answer what was asked?), and completeness (is anything critical missing?).
The three layers form a diagnostic chain: task fails → check the tool-call layer to localize wrong-tool vs. wrong-parameter → tools all correct but task still fails → check the output-quality layer. Blend the three into one score and you'll never know whether to fix the prompt, the tool descriptions, or swap the model. Score them separately. Read the trends separately.
| Layer | Question it answers | How to score | What to fix on failure |
|---|---|---|---|
| Task success | Did the user's job get done? | End-state assertions + LLM judge | Overall flow / task-decomposition prompt |
| Tool-call correctness | Were the tools used right? | Code assertions (deterministic) | Tool descriptions / parameter schemas / tool count |
| Output quality | Is the reply itself sound? | LLM judge + rubric | Style prompt / citation constraints |
3. The Golden Set: Start With 20 Real Tasks
An eval set doesn't need to be big. It needs to be real. The industry rule of thumb is 20–50 real tasks to start — fewer than 20 and statistical noise drowns the signal; more than 50 before you've run a single round and you'll probably abandon maintaining it. Small and runnable every week beats big and aspirational.
"Real" is the operative word: cases must be picked from actual user requests, not invented at your desk. Self-authored cases have a fatal flaw — you'll unconsciously write inputs your agent already handles. Go dig through production logs, support tickets, and problems you hit while dogfooding your own product. Pick three kinds:
- Happy path (~50%): the most common normal tasks. This is your home turf — no iteration may lose points here. A single lost point is an incident.
- Edge cases (~30%): ambiguous input, missing parameters, multi-step flows, requests needing clarification. "Show me last month's books" — does "last month" mean calendar month or trailing 30 days? Should the agent ask or assume? These cases test judgment.
- Refusal / safety cases (~20%): requests it should decline or escalate — out-of-permission actions, abusive users, prompt-injection attempts. Many teams have zero of these in their eval set until the first incident.
The standard structure of one case
Copy this template directly. YAML for humans, JSON for machines, same content:
# golden_case.yaml — golden test set case template
id: invoice-archive-007
category: edge # happy_path | edge | refusal
layer: task # task | tool_call | output_quality (primary layer under test)
input: "File the attached invoice, and tell me the amount"
context: # reproducible preconditions
files: ["invoice_2026-09_acme.pdf"]
today: "2026-10-10" # time-sensitive cases MUST pin "today"
expected_behavior: # expected behavior (for judges and human review)
- Call archive_invoice with month="2026-09" (inferred from invoice content, NOT the current month)
- Reply includes amount "¥12,800" and the filing month
- Must not invent information absent from the invoice
scoring:
- type: code # code scorer
assert: tool_called("archive_invoice") and param("month") == "2026-09"
- type: code
assert: reply_contains("12,800")
- type: llm_judge # LLM judge
rubric: faithfulness # see rubric template in section 4
pass_threshold: 4 # 1-5 scale, >=4 counts as pass
tags: ["invoice", "date-reasoning"]
source: "prod-incident-2026-10-03" # origin: incident id / user report / hand-authored
added_on: "2026-10-10"
Note the today field: every time-sensitive case must hard-code "today", or your eval set rots with time — three months later "last month" points somewhere else, expectations stop matching, and scores drift for no reason. The most commonly overlooked detail in eval-set maintenance.
The production-incident → regression-case pipeline
An eval set is not write-once. Its most valuable growth mechanism is turning every production incident into a case. Fix the process as an SOP:
- Within 24h of the incident: store the user's raw input (sanitized) verbatim as the new case's
input; put the incident ticket insource; - During the postmortem: write up
expected_behavior(what should have happened) and add at least one code-scorer assertion — whichever layer broke gets the assertion; - Verify the fix: run this case first after the fix, then the full golden set — confirm fixing A didn't break B;
- Monthly pruning: delete cases that no longer represent the real distribution (e.g., a deprecated feature). Keep the set small and sharp.
Keep this up for three months and your golden set becomes your most valuable asset: the constitution of your agent's behavior, which every iteration must pass before shipping.
4. Two Kinds of Scorers: Code Assertions + LLM Judges
Scorers come in two kinds with a bright line between them: whatever code can judge, code judges — never an LLM. LLM judges are expensive, slow, and unstable; reserve them for what code cannot judge (open-ended output quality). This division of labor saves you more than half your eval cost and debugging pain.
Code scorers: determinism is a virtue
Tool-call-layer evaluation is almost entirely code's job. The idea is simple: run the agent, take its tool-call trajectory, assert against it item by item. Pseudocode below — translate into your language and it runs:
# Code scorer pseudocode: score one agent run's tool-call trajectory
def score_tool_calls(trajectory, case):
results = {}
calls = trajectory.tool_calls # ordered list: [{name, args, result}]
# 1. Tool selection: expected tool called, forbidden tools not called
results["tool_selection"] = (
expected_tool in [c.name for c in calls]
and not any(c.name in case.forbidden_tools for c in calls)
)
# 2. Parameter correctness: validate key params one by one (type/enum/format)
call = first_call(calls, expected_tool)
results["params"] = all([
isinstance(call.args.get("month"), str),
re.match(r"^\d{4}-\d{2}$", call.args.get("month")), # date format assertion
call.args.get("month") == case.expected_params["month"],
])
# 3. Ordering: dependencies satisfied (lookup before archive)
results["ordering"] = (
index_of(calls, "lookup_invoice") < index_of(calls, "archive_invoice")
)
# 4. Output structure: final reply/artifact matches schema
results["schema"] = jsonschema_validate(
schema=case.output_schema, instance=trajectory.final_output
)
# 5. Keywords/regex: reply contains key facts, contains no banned phrasing
results["keywords"] = (
all(kw in trajectory.reply for kw in case.must_contain)
and not any(bad in trajectory.reply for bad in case.must_not_contain)
)
return results # each item True/False, scored separately — no hasty weighting
The key design decision: score every item separately; don't rush to weight them into one number. Weighting destroys information, and the weights are made up. Let each item's pass rate draw its own trend line first; revisit the single-score question after three months of data.
LLM-as-judge: only where the knife is sharpest
Output quality (faithfulness, relevance) is beyond code's reach — that's where you hire an LLM judge. But "have a model give it a score" is the easiest way to build a useless judge. Unconstrained judge scores are noise. Four constraints, all mandatory:
- Rubric first: the judge doesn't "vibe-score"; it grades item by item against a rubric. Copy the template below;
- Shuffle to kill position bias: when comparing two versions' outputs, randomize A/B order — LLM judges have significant position bias, favoring whichever they see first;
- Separate the judge from the contestant: never let the same model be both athlete and referee. Use a different model family (or at least a different size) as judge; ideally two judges cross-validating;
- Calibrate against human labels: hand-label 30–50 samples first, then measure judge-vs-human agreement. Any judge configuration below 80% agreement is unusable — tune the rubric or swap the judge model until it aligns.
Rubric template (faithfulness dimension, 1–5 scale — reuse by renaming the dimension):
# LLM judge rubric template: faithfulness
You are a strict evaluator. Grade per the rubric below. Output JSON only, no explanations.
Inputs:
- User question: {{user_input}}
- Grounding context (retrieval/tool output): {{grounding_context}}
- Agent's final reply: {{agent_reply}}
Rubric (1-5):
5 = Every factual claim in the reply is grounded in the context; no key information omitted
4 = Claims are essentially grounded; only trivial wording extrapolation
3 = One claim unverifiable from the context, but the core conclusion stands
2 = A claim contradicts the context, or key information is missing
1 = Widespread fabrication, or the answer addresses things the context never supports
Output format:
{"score": <integer 1-5>, "violations": ["<exact sentence violating the rubric>"], "confidence": "<high|medium|low>"}
Pass line: score >= 4
Note: samples with confidence=low are routed to human review and excluded from the pass rate.
That confidence field is the design most teams forget: let the judge say "I'm not sure," and route the unsure ones to humans. It's more honest than swallowing a low-confidence score — and it's how you guard eval credibility on a budget.
On framework choice, one paragraph suffices: DeepEval, RAGAS, TruLens and friends are essentially "these two scorer types, productized, plus a scoring pipeline." When choosing, check three things — how pleasant assertions are to write, how painful CI integration is, and whether the judge model's version can be pinned. Don't learn a framework for the framework's sake; with the method right, 200 lines of hand-rolled script are perfectly enough.
5. The Anti-Self-Deception Checklist: Four Traps and Fixes
The biggest enemy of an eval system isn't technical difficulty — it's self-deception. Four traps below; I've watched real teams fall into every one. Each fix is one sentence:
| # | Trap | Symptom | Fix |
|---|---|---|---|
| 1 | Only testing easy cases | Pass rate sits at 95%+ forever; production still blows up | Enforce the ratio: happy path ≤50%, edge + refusal ≥50%; add 10 real failure samples from production logs every quarter |
| 2 | Only reading the average | Total score rises while a whole category of critical tasks silently fails | Build a category × layer pass-rate matrix; alert when ANY cell drops below threshold — never trust the total |
| 3 | Judge model version unpinned | No code changed, scores drift a little every week, trend lines become meaningless | Pin the judge's model ID + snapshot version in config; changing the judge = re-running human calibration, old scores voided, baseline rebuilt |
| 4 | Thresholds pulled from thin air | "Ship at 80% pass rate" — why 80%? Nobody can say | Derive thresholds from the baseline: run 2 weeks to establish it, threshold = baseline − tolerance band (e.g., 3 points); thresholds may only move UP, never down — lowering one requires a written justification |
Trap 4 deserves one more paragraph. The correct posture for thresholds is "up only, never down": the baseline is your agent's actual current level, and a threshold's job is "no regressions," not "define excellence." When thresholds keep firing and the team wants to lower them "to make CI green" — stop. That usually means a wrong case slipped into the golden set, or the agent genuinely regressed. Write "lowering a threshold requires a written justification" into team norms; that rule alone is worth more than any threshold number.
One more that's easy to miss: the eval set itself rots. The product shipped a redesign, a tool got deprecated, users phrase things differently — but the cases are frozen three months back. Scores keep climbing while measuring a world that no longer exists. Do a quarterly "case freshness" pass: sample 20% of cases against recent production logs and check whether the input distribution still looks like real users.
CI 发布门禁拦住退化的 Agent
6. The CI Gate: Make Evals Able to Block a Merge
Evals that live in a notebook have no teeth. An eval's true form is a CI gate: every PR runs the golden set automatically, and merges below threshold get blocked. Four design points:
1. Keep the golden set small and fast
CI doesn't run your 200-case full suite — it runs a 20–30-case "smoke subset": covering all three layers and all three categories, each case time-boxed (say, 120 seconds, timeout = fail). Goal: finish in 10 minutes. The full suite runs nightly. Slow CI = developers bypassing CI = a dead gate.
2. Gate configuration sketch (pseudo-YAML)
# eval-gate.yaml — CI eval gate sketch (translate to your CI system)
trigger:
on: pull_request
paths: ["prompts/**", "tools/**", "agent_config.yaml"] # only on agent-related changes
eval:
suite: golden_smoke # smoke subset, 20-30 cases
judge_model: "judge-model-v2026-09-snapshot" # pinned version, see section 5
timeout_per_case: 120s
max_cost_usd: 5.00 # cost cap per eval run; exceeding it = fail
gates: # per-layer gates, no weighting
- layer: tool_call
metric: pass_rate
threshold: 0.97 # strictest on tool calls: zero tolerance for regression
- layer: task
metric: pass_rate
threshold: 0.90 # baseline 0.93, 3-point tolerance band
- layer: output_quality
metric: pass_rate
threshold: 0.85
on_failure:
- block_merge: true
- comment_on_pr: true # post the failing case list as a PR comment
- artifact: full_trajectory # keep complete call trajectories for debugging
3. Per-layer thresholds, no total-score threshold
Echoing section 2: 97% on tool calls, 90% on tasks, 85% on output quality — each blocks independently. Why strictest on tool calls? Because it's deterministic — a regression in something deterministic means the code or config is genuinely broken. There's no "the judge was in a bad mood" excuse.
4. Cost and latency are metrics too
Don't let evals themselves get expensive: cap per-run cost (say, $5), cache judge calls (cache identical inputs' judge results for 7 days); and put the agent's average token consumption and end-to-end latency into the gate as well — an iteration with unchanged quality that's 3× slower and 5× pricier shouldn't ship silently either.
7. The Production Loop: Evals Are Ongoing Engineering
The offline golden set answers "don't regress before shipping," but it has a natural blind spot: it tests yesterday's users. The real world moves, so the eval system needs a production loop feeding back in. Three practices, in ascending order of investment:
1. Sample real traffic and score it with the same scorers
Sample production traffic daily (say, 1%), sanitize it, run the same code scorers + LLM judges from CI, and push scores to a dashboard. This is your agent's true-capability thermometer. Offline at 95%, sampled production at 78% — that gap is itself the most important signal: your golden set has drifted from the real distribution.
2. Shadow runs: let the new version rehearse before performing
Before a major release, run the new agent version in shadow mode against live requests: execute for real, but return nothing to the user — only record trajectories and scores. Run it for a week, compare score distributions against the production version. Shadow runs are the cheapest way to catch "fixed A, broke B" — real traffic, zero user consequences.
3. Weekly review: feed failure samples back into the test set
Fixed cadence, 30 minutes a week, three things on the agenda:
- The 10 lowest-scoring samples from this week's production sampling — true failure or scorer misfire? Misfire → fix the scorer; true failure → immortalize as a regression case per the section-3 SOP;
- The offline-vs-production gap trend — a widening gap means the golden set needs new cases;
- A judge-agreement spot check — hand-review 10 LLM-judged samples weekly; agreement below 80% triggers recalibration.
Once this loop spins up, you get compounding returns: the golden set looks more like the real world, evals get more trustworthy, iterations get bolder. Without the loop, an eval set becomes correct nonsense within three months — beautiful scores measuring users who don't exist.
Closing: Build the Scaffold First, Optimize Later
A minimum starter checklist — one afternoon gets you to v1:
- Pull 20 real tasks from production logs (10 happy + 6 edge + 4 refusal) and write them as YAML per the section-3 template;
- Write 10 code assertions (tool selection, parameters, ordering, schema, keywords) — get the tool-call layer scoring first;
- Copy the section-4 rubric template, run a 30-sample human calibration, push judge agreement past 80%;
- Wire the smoke subset into CI with three layer thresholds — and let it block one merge. The first blocked merge teaches the team why this exists better than any doc;
- Book the weekly 30-minute review on the calendar first, then talk about optimization.
Don't chase 200 cases, fully automated judges, and a perfect dashboard on day one. An eval system's value isn't in how complete it is — it's in the fact that it exists and runs every week. A 20-case set that runs weekly and has blocked one merge beats a 200-case "perfect plan" gathering dust.
Stop iterating on vibes. Start with these 20 cases, today.
Related articles

You don't need a support team — you need a support system. This guide walks solo developers through the full playbook: ticket triage, a three-layer defense funnel, ticket-deflecting FAQs, an AI draft pipeline with copy-paste prompt templates, five canned-response templates, automation red lines, and a weekly 30-minute review SOP. Every section ships with templates you can use today.

Simon Willison shipped a Newsletters index page for his blog almost entirely by voice — chatting with the Codex tab in the ChatGPT desktop app while cooking dinner in his kitchen, barely touching the keyboard. When a top-tier practitioner starts coding with his mouth, voice + agents stop being a gimmick and become real productivity. A teardown of his playbook, where this workflow breaks, and the minimal setup to copy him.

On October 8, 2026, Anthropic put Claude Dashboards and Claude Motion into beta: dashboards built from plain-language questions on live company data, and animations generated as editable code rather than video-model footage. Docs, Slides, and Design went GA on all plans, with 45M+ artifacts created to date.