Stop Shipping AI Features on Gut Feeling: A Practical Evals System for Vibe-Coded Projects
Without evals, every model swap is a gamble. This guide gives solo teams an evaluation system built in half a day: mining 50–200 exam questions from real logs, rule assertions vs LLM-as-judge (and taming judge bias), a regression pipeline, the production sampling flywheel, and launch gates with rollback — plus copy-paste run.py, CI configs, and checklists.

The verdict up front
Without evals, every model swap is a gamble.
That's not a metaphor. "Gamble" means this: you put something unobservable, unreproducible, and silently mutable by its vendor into the core path of your product, then ship it because "I tried it a few times and it seemed fine." In traditional software that process is called "having no tests." In AI features it's somehow been normalized — because model output reads so much like human language that you mistake it for something as stable as code.
It isn't. In 2023, three researchers at Stanford and UC Berkeley (Chen, Zaharia, Zou) tracked how GPT-3.5 and GPT-4 behaved between March and June 2023, and the results were brutal: GPT-4's accuracy at identifying prime numbers was 97.6% in March; by June, on the same questions, from the same "GPT-4," it had fallen to 2.4%. Code generation was worse: the share of directly runnable code dropped from 52% to 10%, because the newer model started wrapping code in extra quotation marks. The vendor issued no announcement, changed no version number, and your prompt hadn't changed a single character — the output had quietly become a different person.
This guide skips the philosophy and hands you something a solo team can build in half a day: a fixed set of exam questions, graders, regression runs, and a launch gate. One goal — so that next time you swap a model or edit a prompt, you're holding cards, not just hope.
1. "I tried it a few times and it seemed fine": the most expensive QA process in your project
Let's price out shipping on gut feeling. It's expensive because of three illusions, each perfectly shaped to fit the vibe-coding comfort zone.
Illusion one: sampling illusion
The 5 examples you tried are all the happy paths you know best. You had exactly those 5 phrasings in mind when you wrote the prompt, so of course the model aces them. Users will hit you with the 6th posture: typos, dialect, half-sentences, two requests crammed into one. You didn't test "how users will ask" — you tested "how you wish users would ask." The gap between the two is the shape of your next production incident.
Illusion two: the author filter
You hand-tuned the prompt, so you grade its output with a filter. A missing key fact and your brain auto-completes "users will figure it out"; a slightly off tone and you think "nobody will care." You're not evaluating the output — you're making excuses for your own work. Letting the author be the judge is a permanent perfect score. That's not a character flaw, it's a structural one.
Illusion three: the time illusion
This is the killer. The "seems fine" you measured today was measured against today's model. Vendor models update continuously: silent fine-tunes, alignment adjustments, even full version replacements — sometimes without telling you. The Stanford study above is the smoking gun: three months, same service, prime-number accuracy from 97.6% to 2.4%. Your "fine" today may expire next month, and you have no mechanism that would notice the expiry.
Two more gambling scenes vibe developers play almost weekly:
- Swapping in a cheaper model to save money. Replace the flagship with a mini, or a vendor three times cheaper, try ten queries, call it "close enough," and cut over 100%. Three months later users complain "the AI got dumber," and you can't even say whether it really got dumber or users just got pickier.
- Fixing a bug with a prompt edit that introduces regressions. Users report scenario A answers poorly, you add instructions to the prompt, scenario A improves, scenarios B, C, D quietly break. You fixed the one in front of you, broke three behind you, and feel like you're "iterating fast."
A solo team's full accounting: one production incident costs firefighting time + lost user trust + your blood pressure at 3 AM rolling back. A minimal eval setup costs half a day plus pennies of API fees per run. The math isn't hard; the hard part is admitting "I tried it a few times" was never testing.
Compress countless late nights into one screenplay and every incident has the same shape: Monday, you swap the flagship model for a lite version three times cheaper, hand-test 10 queries, all pass, cut over 100% that night. Wednesday, a user says "the AI started answering irrelevantly" — you check the logs and find the lite model drops mid-context instructions on long inputs, and all 10 of your test queries were short. Friday, you try to roll back and discover you used a floating alias, so you can't even say which version is actually serving. Sunday, you apologize in the user group and promise "next time we'll test thoroughly." There is no next time — the "next time" without evals uses the same method as this time.
There is no villain in this screenplay. The vendor did nothing wrong (they keep iterating), and you weren't lazy (you really did test 10 queries). The method is wrong: using 10 happy paths to validate a change in a distribution, using one-off hand-feel to replace repeatable measurement. The speed dividend of vibe coding was always earned this way; once the dividend is collected, the bill comes due.
2. What evals are — and aren't
Let's nail the definition first, so we don't talk past each other later.
Evals = a fixed set of inputs + graders + a repeatable scoring run. It doesn't answer "is this output good." It answers "did this version get better or worse than the last one, and by how much." The output is a number, and the change in that number.
With the definition set, here are three "aren'ts" — because these three confusions are the biggest reason evals never get adopted.
Not unit tests
Unit tests assert code logic: deterministic input-output, identical results across a million runs. Evals assert the distribution of model behavior: stochastic, with run-to-run jitter. Your all-green pytest suite has no relationship whatsoever with "did the AI feature get better" — it tests the code around the model, never what the model says. Plenty of teams ship with hundreds of green tests and watch the AI feature explode on launch day: 100% coverage of the deterministic code, 0% coverage of the nondeterministic output. The classic case: you build an endpoint that turns free-text messages into structured tickets; unit tests assert "returns valid JSON with all fields" — all green. One day a colleague tweaks the system prompt for tone, and the model starts inventing order IDs for messages that have none. The JSON is still valid, the fields are still complete, every unit test stays green, and only users complain that "the order IDs in your system are fake." That's the hole evals fill: tests around the code can never see what the model said.
Not human spot checks
Spot checks answer "is this output okay" — a point. Evals answer "did this version get better or worse overall" — a surface. The problem with spot checks is they're unrepeatable and incomparable: you sample 20 this week and feel good, sample 20 next week and feel bad — did the model change, or did your sample change? You can't say. Evals re-run the same exam, so there's exactly one variable: the version under test.
Not public benchmarks
Leaderboards like SWE-bench or MMLU prove "this model is strong." They don't prove "your prompt + this model + your business data" is strong. Benchmarks are talent shows; evals are physical exams — and the exam has to use your body, your users' real inputs. Shipping on a chart-topping model with an untested prompt is like letting someone with a clean bill of health take your exam for you.
One table for the division of labor:
| Unit tests | Evals | Human spot checks | |
|---|---|---|---|
| Tests | Whether the code logic is right | Whether model output behavior improved/regressed | Case quality, experience details |
| Runs | Every commit | Model swaps, prompt edits, scheduled regressions | Fixed weekly slot |
| Answers | Is the code broken | Is this version better or worse than the last | What users actually feel |
| Costs | Near zero | Pennies to dollars of API fees per run | Your time — the most expensive |
| Missing it means | Every code change can crash | Every model swap is a gamble | Experience rots slowly, unnoticed |
My position is blunt: you need all three, but for AI features, evals is the missing plank. You're already writing unit tests, you occasionally spot-check, and evals is probably at zero — while model behavior is exactly the fastest-changing, least controllable part of your product.
How evals die after you build them
Three preventive shots, because eval systems have three classic ways to die — all of them after "we built it":
- Death by "ran once and shelved." The question bank never gets updated again; three months later it's testing last quarter's business. Question banks rot like code — without a steady diet of production failures, it's a souvenir.
- Death by score inflation. To keep CI green, failing questions get deleted, rubrics get watered down, tolerance bands get widened. The score is back to 95 and nothing got fixed. This is worse than no evals — without evals you at least know you're gambling; inflated scores let you believe you're not.
- Death by "nobody reads the report." The pipeline runs daily, reports generate daily, regressions go unaddressed. On a solo team, "nobody" is you — so put "read the report" in the release checklist instead of relying on willpower.
Three sentences for survival: the bank gains new questions monthly, deleting questions needs review, the report gets read before every release. Do those three and your evals survive past month three; past month three, they become something you can't live without.
3. The golden dataset: a 4-step method to mine 50–200 "exam questions" from real user logs
The dataset is 90% of the work in evals, and 90% of the value. The prettiest graders on earth produce garbage scores from garbage questions. On "where do questions come from," I have exactly one strong opinion:
Questions must be mined from real user logs — never invented by you. Questions you write are ones you already know how to answer; questions users write are the ones you don't. Writing your own exam, taking it yourself, and grading it yourself is a permanent perfect score — the most common self-deception in evals practice.
How, concretely? Four steps.
Step 1: Mine — export real inputs from production logs
Dig through the last one to three months of production logs and export what users actually sent to the AI feature. Anonymize first: names, phone numbers, order IDs, addresses all become placeholders (e.g. [NAME], [ORDER_ID]). Write the anonymization script once; reuse it every time you mine.
Mine with intent to cover three input types — miss one and your question bank has a blind spot:
- High-frequency inputs (~50%). The most common phrasings. This is your bread and butter; losing points here means a broad outage.
- Failure inputs (~30%). The ones where users clicked "not helpful," hit regenerate, or escalated to a human. These are lessons already paid for in tuition — leaving them out of the bank means the tuition was wasted.
- Long-tail inputs (~20%). Randomly sample the odd, obscure ones. The long tail is where models embarrass themselves most, and it's exactly what gut-feel testing never covers.
Step 2: Sort — dedupe, cluster, label
Raw logs are 80% repeated phrasings. Dedupe first: simple text-similarity dedup works; embedding clustering is the fancier option. Then label along two axes:
- category (task type): e.g. refund questions, how-to questions, troubleshooting, chit-chat. Keep at least 5 per category, or a whole category could collapse without you noticing.
- difficulty: easy / medium / nasty. The "nasty" tier is for emotional, information-poor, deliberately tricky inputs — the most painful production cases all live here.
Step 3: Label — write "pass conditions," not "model answers"
This is the most important step and the easiest to get wrong. For open-ended output, don't write a full "model answer" — when the model phrases a correct answer differently and you mark it wrong, your evals become a memorization contest. Write pass conditions instead, in three parts:
- must_contain: what has to appear (keywords, key facts).
- must_not_contain: what must never appear (invented order IDs, promises outside policy, profanity).
- judge_rubric: for the parts needing semantic judgment — the grading standard for an LLM judge or a human, one sentence on "what passes, what fails."
The data format looks like this — one JSONL line per question:
{"id": "refund-001", "category": "refund", "difficulty": "medium",
"input": "I bought the membership three days ago, can I get a refund?",
"expected": {"must_contain": ["7-day", "refund"],
"must_not_contain": ["non-refundable"],
"judge_rubric": "Must cite the 7-day no-questions-asked refund policy; must not invent refund links or support phone numbers"}}
{"id": "refund-002", "category": "refund", "difficulty": "nasty",
"input": "This useless membership is trash, refund me NOW or I'm filing a complaint!",
"expected": {"must_contain": ["refund"],
"must_not_contain": ["regulator", "threat"],
"judge_rubric": "Tone must be calming but not servile; no compensation beyond policy; no arguing with the user"}}
{"id": "usage-014", "category": "how-to", "difficulty": "easy",
"input": "How do I export my chat history?",
"expected": {"must_contain": ["Settings", "export"],
"must_not_contain": [],
"judge_rubric": "Steps must match the current UI; must not describe retired entry points"}}
Step 4: Lock — commit to git, version it, append-only
The question bank is a code asset; version it. Two iron rules:
- Questions may only be added, never deleted. Deleting a question deletes history and covers up regressions. A stale question (retired feature) gets marked deprecated, not deleted — changing the bank itself needs review.
- Every production incident review starts with one question: is this in the question bank? If not, the review isn't over. That's the only reliable mechanism for the bank to keep growing.
On quantity: 50 is the starting line — a run takes minutes and pennies, so you can afford to run it on every prompt change. 200 is the comfort zone. A solo team shouldn't start at 500: banks rot, business changes, old questions need maintenance. Get 50 running first, then grow. Remember the oft-validated wisdom: a tight 50-question bank that runs on every change beats a 1,000-question bank that runs once and rots.
One anti-pattern to close: bulk-generating questions with AI, then testing and grading with AI. AI-written questions are what the AI thinks users ask; the standard is what it thinks is right; the grader is itself — a closed loop, a permanent perfect score, pure self-soothing. The one job AI may do: rephrase and expand already human-labeled questions into variants (3 phrasings of the same test point). Writing questions and setting standards is human work.
4. Grader design: rule assertions vs LLM-as-judge
Questions ready. Next: who grades? Two approaches, each with its own temperament. My position first:
If a rule can grade it, don't hire an LLM judge for a single question. Judges are luxuries: expensive, slow, and biased. Rule assertions run in milliseconds, cost nothing, and never jitter — they're the ballast of evals.
The rule-assertion toolbox
Don't underestimate rules — they cover more than you'd think:
| Scenario | Assertion | Example |
|---|---|---|
| Classification / extraction | Exact or set match | Intent label must equal expected value |
| Structured output | JSON Schema validation | All fields present, correct types, no extras |
| Key information | must_contain / must_not_contain | Refund-policy keywords appear; invented order IDs don't |
| Numbers | Range assertion | Confidence between 0–1; price positive |
| Format | Regex | Dates, phone numbers, markdown table headers |
| Semantic closeness | Embedding similarity ≥ threshold | Paraphrase tasks; 0.8 is a common starting point (tune with human-labeled data) |
| Length | Interval assertion | Summary under 200 chars; title 10–30 chars |
The decision rule in one sentence: whenever "writing the rule" is cheaper than "hiring the judge," pick the rule, no contest. There's exactly one situation that requires a judge — open-ended answer quality: is the tone appropriate, is the explanation logical, is the soothing copy on point, is the creative writing any good. These cost more to encode as rules than to judge, and only then does the judge take the stage.
A judge's four biases — never mistake it for "objective"
LLM-as-judge means calling a second model as referee, scoring outputs against your rubric. Useful, but with four empirically documented systematic biases:
- Verbosity bias: longer answers get free points. Between two equally good answers, the judge probably picks the longer one.
- Position bias: in pairwise comparison, the answer seen first has the edge. Swap the order and the winner can flip.
- Self-preference: judges score their own model family's outputs higher. A GPT judge grading GPT vs Claude outputs starts with a tilted scale.
- Tolerance drift: the same judge is strict today, lenient tomorrow. So the judge's model version must be pinned — change the judge, and the entire score history is void.
Six moves to keep bias on a leash
Bias can't be eliminated, only managed. Six moves, in order of importance:
- Binary rubrics, not 1–10 scores. "PASS / FAIL" is an order of magnitude more stable than "7 or 8." Precision past the decimal point is hallucination.
- Write "what fails" before "what passes." The more specific the failure conditions, the more honest the judge. Vague rubrics breed vague scores.
- Pin the judge's model version. A concrete version ID, never a floating alias like latest. Change the judge model, re-run all history.
- Blind grading. Don't tell the judge whose output it's grading. In A/B comparisons, hide candidate identities — just "Answer A / Answer B."
- Swap order and run pairwise comparisons twice. AB order and BA order, one run each; it only counts as a win if it wins both. A single-run comparison isn't trustworthy.
- Calibrate against human labels before the judge ships. Sample 30–50, humans grade first, judge grades second, compute agreement. Below 85%, the judge doesn't ship — fix the rubric first, or swap the judge model. An uncalibrated judge's scores are just noise shaped like numbers.
A judge prompt template you can copy (swap in your own rubric):
You are a grader, not an advisor. Output only PASS or FAIL plus a one-line reason.
1. The answer cites the 7-day no-questions-asked refund policy
2. No invented refund links or phone numbers outside the policy
1. Promises compensation beyond the policy
2. Argues with the user or uses threatening language
3. Doesn't answer the question
Answer under review:
---
{output}
---
Check FAIL conditions one by one first, then PASS conditions. Output only:
verdict: PASS / FAIL
reason: one line
A final reference ratio: the community-validated rule of thumb is 60% rule assertions + 30% LLM judge + 10% human review. Not iron law — a starting point. Your v1 can absolutely be 90% rules + 10% human, and only hire the judge when open-ended cases back you into a corner. The later you hire it, the better your rule-encoding is.
Appendix: how to set the embedding-similarity threshold
"0.8 is a common starting point" — but a starting point isn't an answer. Thresholds must be tuned on human-labeled data. The method is unglamorous but works:
- Pull 50 pairs (model output, reference paraphrase) from the bank; humans first judge "same meaning / different."
- Compute embedding similarity per pair, sort by score.
- Find the cut point minimizing disagreements: above threshold = "same," below = "different"; count conflicts with human judgment.
- Set the threshold at the minimum-conflict point, then tighten by 0.02 — better a false kill (human re-review) than a false pass (bad output slipping through).
Thresholds are bound to the embedding model: change the embedding model, re-tune the threshold. Same rule as "recalibrate when the judge model changes": any model inside a grader whose mental version changed voids the scores.
5. Regression evals: the minimal pipeline for before/after comparison on model or prompt changes
Questions and graders ready — the core move of evals is just one, boring to the point of comedy:
Run once before the change, save it as the baseline; run again after; diff. A score without a baseline is theater — the number itself means nothing, the change means everything. Is 92 good or bad? No idea. But 92 → 88 is definitely bad news.
What a regression report looks like
A usable regression report has five parts, none optional:
- Total score: baseline 92.0 → this run 91.5 (Δ -0.5). Obvious at a glance.
- Per-category scores: refunds 95→94, troubleshooting 88→82 — the total only slipped 0.5, but "troubleshooting" fell 6 points, and that's the one that kills. Totals lie; category scores don't.
- Regressed case list: exactly which questions failed, in which category, with output snippets. All debugging starts here — without it, the report is correctly-worded nonsense.
- Cost and latency: how much the cheaper model saved, how many milliseconds it added. Accuracy unchanged but latency doubled is still a regression — users don't pay for your cost optimization.
- Environment info: model version, prompt version, run timestamp. Without these three fields, three months from now you'll have no idea what this report measured.
What it looks like in the terminal — conclusions at a glance, no more "feeling":
$ python evals/run.py --compare evals/baseline.json
{
"score": 0.880, "total": 120,
"by_category": {"refund": 0.95, "how-to": 0.93, "troubleshooting": 0.78, "chit-chat": 0.90},
"failed": ["trouble-017", "trouble-023", "trouble-031"],
"cost_usd": 0.42, "p95_latency_ms": 1830
}
baseline=0.920 now=0.880 delta=-0.040
REGRESSION DETECTED, failing the build.
Regressed categories: ['troubleshooting']
Seeing "troubleshooting 0.88 → 0.78" tells you to go read trouble-017 and friends instead of staring at the total. Totals are for outsiders; category scores and the failed-case list are for engineers.
The minimal pipeline: three pieces
Regression evals need no platform — three pieces suffice:
- run.py: reads golden.jsonl → calls the model → runs graders → writes results.json + terminal report → compares against baseline.json → exits nonzero on regression. Full copy-paste template in section 8.
- baseline.json: the score snapshot of the last release, in git. After every real release, re-run with that release's code and model and update the baseline — the baseline must always correspond to "what's actually serving," or the comparison is meaningless.
- CI job: auto-triggers on prompt edits, model swaps, or AI-pipeline code changes. Regression → pipeline goes red → merge blocked. YAML template in section 8.
One word on trigger frequency: don't run on every push — evals burn API money. Run only when prompts, the question bank, or AI-pipeline code change, plus one scheduled full run nightly. That usually costs a few coffees a month — two orders of magnitude cheaper than a 3 AM firefight.
Cold start: how to set the first baseline
"I don't even have a first score — what do I baseline against?" The most common starting blocker. Answer: run today's serving version once, right now — that is your baseline. Don't agonize over whether it's high or low; a baseline's job isn't "looking good," it's "being honest." Steps:
- Freeze the current serving version (model version + prompt version), tag it.
- Run the full question bank on it, store the score in baseline.json untouched — no manual "optimization."
- If the score is shockingly low (say, 70), don't rush to tune the prompt — check whether your questions were labeled too strictly. Raising the score comes later, after the baseline is set.
An unwritten rule of first baselines: it will make you uncomfortable. Hand-labeled nasty questions always fail a few on the first run. That's good — the first deliverable of evals is an honest list of "where we actually suck." Without that list, every "seems fine" before it was illusion.
The gate rule: block "regression," not "imperfection"
This is where eval systems get most misused. Newcomers stand up evals and immediately set a high bar: "must score 95 to ship." Wrong. Your feature scores 92 today; demanding 95 means two outcomes: nobody dares ship, or everyone starts watering down questions and deleting hard ones. Both are disasters.
The correct gate philosophy is one sentence: scores may only go up, never down. Baseline 92, tolerance band 3%, gate at 89. Scored 91.5? Ship — that's noise. 88? Stop, find which category regressed. The gate watches the direction of change, not the absolute height. Absolute height is for quarterly planning, not release gates.
A statistics footnote: don't hold meetings over 1% jitter
LLM output is stochastic; single runs jitter. Three engineering habits iron it out:
- Run each question 3 times, take the majority/median. 2-of-3 passes counts as pass. Don't conclude from 1 run, and don't burn money on 10.
- Set the tolerance band at 2–3%. Treat movement inside the band as noise — no meetings, no rollbacks, no prompt edits. Chasing noise will make you break a prompt that was fine.
- For big changes, run 5 times and look at the distribution. Model or vendor swaps are the bone-breaking kind — run several times to confirm stability, don't let one lucky run fool you.
Remember: evals are a statistical creature, not a deterministic machine. Treating evals like unit tests ("must pass 100%, one jitter is a bug") will make you abandon them in week one. The right use is a weather forecast: watch trends and changes, don't obsess over a single degree.
6. The online closed loop: production sampling + human spot checks + an ever-growing question bank
Offline runs only guarantee "the questions you've tested don't regress" — but users will forever invent questions you never tested. The second half of evals lives in production: turning real-world failures into questions, continuously.
Production sampling: three buckets, not random scoops
Sample production logs daily (or weekly, depending on volume), targeting three buckets:
- User downvotes / "regenerate" clicks: always sample. These are user-hand-labeled failure scenes — the highest-value samples you own.
- Retries, timeouts, human handoffs: always sample. Pipeline-level failures usually mirror model-output failures.
- Random 1%: guards against survivorship bias. Looking only at failures convinces you the world is all failures; the random sample tells you whether the bread and butter still holds.
Make sampling a scheduled job, never a manual chore. Pseudo-SQL you can adapt to your log schema:
-- Runs daily at small hours: pull three buckets into eval_sampling
-- 1) user downvotes / regenerate clicks
SELECT input, output FROM ai_logs
WHERE created_at > NOW() - INTERVAL '1 day'
AND (user_feedback = 'bad' OR regenerated = true)
LIMIT 50;
-- 2) retries / timeouts / human handoffs
SELECT input, output FROM ai_logs
WHERE created_at > NOW() - INTERVAL '1 day'
AND (retry_count > 0 OR timed_out OR handed_to_human)
LIMIT 50;
-- 3) random 1% (hash the primary key for repeatable sampling)
SELECT input, output FROM ai_logs
WHERE created_at > NOW() - INTERVAL '1 day'
AND MOD(HASH(id), 100) = 0
LIMIT 100;
The spot-check protocol: 30 minutes a week, not a minute more
Spot checks die from unsustainability, not inaccuracy. So the protocol must be too light to fail:
- Same time every week, 20 samples (from the sampling buckets).
- Each gets pass / fail plus a one-line reason. No essays.
- For failures, decide on the spot: into the question bank, or onto the todo list.
20 samples, 30 minutes, one coffee. A cadence you can keep for a year keeps evals alive for a year. A 100-sample, 2-hour plan dies in week three — including for you.
The flywheel: turn tuition paid into receipts
The full flywheel:
- A production failure (user complaint / downvote / something you spotted).
- A 5-minute retro: prompt problem, model problem, or out-of-scope input? Just these three buckets, no essays.
- Write it as a question, commit it to golden.jsonl with anonymized input and pass conditions.
- The next similar change gets blocked by the regression suite before release.
The question bank is the receipt for every tuition payment you've made. Lose the receipts and the tuition was wasted. Two health metrics, glanced at monthly:
- New questions per month: two straight months of zero means the retro process died, or nobody's looking at samples.
- Share of production incidents covered by the bank: target 80% of incidents having a sibling question in the bank. Miss it and your "failure → question" pipeline is broken.
7. The launch gate: what score is shippable
Everything so far was "how to measure." This section is "what to do with the measurement." Evals without a gate is a physical exam you never read — full ceremony, zero effect.
How to set the gate line
Don't pull "must score 95" out of thin air. Gate line = baseline − tolerance band. Baseline 92, band 3%, gate at 89. Below 89, the release stops. The philosophy: block "regression," not "imperfection." Your feature scores 92 today; demanding 95 means nobody ships, or everyone waters down the questions.
New features and major prompt rewrites get their own lines — looser at first (say, baseline − 5%), tightened after two stable weeks. The gate line itself gets versioned; every adjustment leaves a record and a reason. Keep it in evals/gate.md:
# Launch gate
- Current baseline: 0.920 (2026-10-11, model gpt-4o-mini-2024-07-18, prompt v14)
- Gate line: 0.890 (= baseline - 0.03)
- Per-category red line: any category dropping > 5% blocks independently
- Change log:
- 2026-10-11: v1, baseline 0.920, gate 0.890 (reason: first baseline)
- 2026-10-25: baseline raised to 0.935, gate to 0.905 (reason: prompt v15 improved troubleshooting +4%)
Note the second entry: when the baseline rises, the gate rises with it. A gate that only ever lowers becomes decoration — "89 is easy anyway." The gate is alive; it should climb with your quality waterline.
Scores dropped — who's responsible
One sentence, no exceptions: whoever edits the prompt or swaps the model runs the evals and pastes the report into the PR.
Solo teams don't skip this either — it's not for show, it's for you in three months. You will forget; you'll forget "why 88 seemed shippable at the time." The report pasted in the PR is evidence for your future self. Add one line to the PR template:
## Evals regression report
- baseline: 0.920 (2026-10-11)
- this run: 0.915 (Δ -0.005)
- category changes: no category down more than 2%
- verdict: no regression, shippable / REGRESSED, cause: ___
Scores dropped — how to roll back
The gate blocked the merge, but sometimes "must ship tonight" is real. That's what the rollback trio is for — not courage:
- Pin the model version hard. Write
gpt-4o-mini-2024-07-18, never a floating alias likegpt-4o. Floating aliases are the vendor's backdoor into your system: they change the pointer one day and your "rollback" rolls back nothing. - Version the prompt. Prompts are code — in git, tagged together with the model version and the question bank. Rolling back half (prompt without model) is rolling back nothing.
- A feature flag that cuts back to the old path in one flip. New path scores badly? Flip the flag, traffic returns to the previous version. This switch must have been genuinely rehearsed once — an unrehearsed switch will jam exactly when it matters.
The exception path: when you truly must ship it hard
The gate blocked the merge but the business says "tonight, no matter what." Leave exceptions a documented road, or people learn to route around the gate:
- Ship behind the feature flag at 5% traffic, watch sampling and user feedback, decide on full rollout after 24 hours.
- The exception note records three things: why this couldn't wait for the score to be fixed (business reason), whose call it was (a name), and the rollback condition (flip back if the score drops X more).
- The note goes into git, next to the eval report. At the next retro it's the evidence of "why we took the risk."
An undocumented exception is "breaking the rules"; a documented one is "a risk decision." A solo team's biggest risk isn't slowness — it's nobody remembering why that decision was made three months later. And that nobody is you.
The gate's meaning, said plainly at last: it's not about blocking every bad release — impossible, the bank will always have blind spots. It's about every release having a record, an owner, and a rollback button. Bad releases aren't scary. What's scary is not knowing how it broke and not being able to unbreak it.
8. A solo team's minimal evals setup: built in half a day
Theory done — hands on. This setup has one goal: built this afternoon, usable tonight. No evals platform, no dashboards, no spending — one working run.py beats ten slides of "evaluation system planning."
Directory layout
evals/
├── golden.jsonl # question bank: one JSONL line each, with pass conditions
├── graders.py # graders: rule-assertion functions
├── run.py # runner: runs questions, grades, diffs baseline, prints report
├── baseline.json # score snapshot of the last release (in git)
├── judge_prompt.txt # LLM judge rubric template (only if you have open-ended questions)
└── requirements.txt # openai / anthropic, pick one, plus your main SDK
graders.py: rule-grader template
import json, re
def grade(case, output):
"""Returns (passed: bool, reasons: list[str])"""
exp = case["expected"]
reasons = []
ok = True
for kw in exp.get("must_contain", []):
if kw not in output:
ok = False; reasons.append(f"missing keyword: {kw}")
for kw in exp.get("must_not_contain", []):
if kw in output:
ok = False; reasons.append(f"forbidden term present: {kw}")
return ok, reasons
def grade_json_schema(case, output, required_fields):
"""Structured output: valid JSON first, then complete fields"""
try:
data = json.loads(output)
except Exception:
return False, ["output is not valid JSON"]
missing = [f for f in required_fields if f not in data]
if missing:
return False, [f"missing fields: {missing}"]
return True, []
def grade_length(case, output, min_len=0, max_len=10_000):
n = len(output)
if not (min_len <= n <= max_len):
return False, [f"length {n} outside [{min_len}, {max_len}]"]
return True, []
run.py: run + diff-against-baseline template
import json, sys, statistics
from graders import grade
MODEL = "gpt-4o-mini-2024-07-18" # pin the exact version, never an alias
TOLERANCE = 0.03 # tolerance band: 3%
REPEATS = 3 # 3 runs per question, median wins out jitter
def call_model(prompt_input):
# TODO: swap in your real call: build prompt, call model, return text
raise NotImplementedError
def run_once(cases):
results = []
for c in cases:
outs = [call_model(c["input"]) for _ in range(REPEATS)]
passes = [grade(c, o)[0] for o in outs]
# median of 3: at least 2 passes counts as pass
passed = statistics.median(passes) == 1
results.append({"id": c["id"], "category": c["category"],
"passed": passed})
return results
def summarize(results):
total = len(results)
score = sum(r["passed"] for r in results) / total
by_cat = {}
for r in results:
by_cat.setdefault(r["category"], []).append(r["passed"])
by_cat = {k: sum(v)/len(v) for k, v in by_cat.items()}
failed = [r["id"] for r in results if not r["passed"]]
return {"score": round(score, 3), "by_category": by_cat,
"failed": failed, "total": total}
if __name__ == "__main__":
cases = [json.loads(l) for l in open("golden.jsonl", encoding="utf-8")]
summary = summarize(run_once(cases))
print(json.dumps(summary, ensure_ascii=False, indent=2))
json.dump(summary, open("results.json", "w", encoding="utf-8"),
ensure_ascii=False, indent=2)
base = json.load(open("baseline.json", encoding="utf-8"))
delta = summary["score"] - base["score"]
print(f"baseline={base['score']} now={summary['score']} delta={delta:+.3f}")
regressed = [c for c, s in summary["by_category"].items()
if s < base["by_category"].get(c, 1) - TOLERANCE]
if summary["score"] < base["score"] - TOLERANCE or regressed:
print("REGRESSION DETECTED, failing the build.")
if regressed:
print("Regressed categories:", regressed)
sys.exit(1)
print("OK: no regression.")
Two details worth noting: 3 runs per question with the median because single runs jitter — don't hold meetings over 1% noise; and the exit(1) is for CI — regression turns the pipeline red, the only "integration" a gate needs.
Have open-ended questions? Add a judge extension point to run.py — rules first, judge as backstop:
def grade_with_judge(case, output, judge_model="gpt-4o-2024-08-06"):
# rules first: if rules fail, fail fast without spending judge money
ok, reasons = grade(case, output)
if not ok:
return False, reasons
rubric = case["expected"].get("judge_rubric")
if not rubric:
return True, [] # no rubric: rule pass = pass
prompt = open("judge_prompt.txt", encoding="utf-8").read()
verdict = call_judge(judge_model, prompt.format(
rubric=rubric, output=output)) # TODO: swap in your judge call
passed = verdict.strip().upper().startswith("PASS")
return passed, [] if passed else ["judge failed: " + verdict[:200]]
The ordering is deliberate: rules are the sieve, the judge is the microscope. The sieve catches the obvious failures; the microscope only examines what the sieve can't decide. Reverse it (judge first, rules after) and every dollar you spend pays for problems a regex could have solved.
CI wiring: GitHub Actions template
name: evals
on:
pull_request:
paths: ['prompts/**', 'evals/**', 'src/ai/**']
jobs:
evals:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with: { python-version: '3.11' }
- run: pip install -r evals/requirements.txt
- run: python evals/run.py
working-directory: evals
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
Don't want to write code? The promptfoo one-liner
The open-source tool promptfoo exists for exactly this: questions and assertions in YAML, one command to run. Minimal config:
# promptfooconfig.yaml
prompts:
- file://prompt.txt # your system prompt
providers:
- openai:gpt-4o-mini-2024-07-18
tests:
- vars: { input: "I bought the membership three days ago, can I get a refund?" }
assert:
- type: icontains
value: "7-day"
- type: not-icontains
value: "non-refundable"
- vars: { input: "How do I export my chat history?" }
assert:
- type: icontains
value: "export"
npx promptfoo@latest eval # run
npx promptfoo@latest view # report dashboard
promptfoo is for "get running first" — migrate to your own run.py later when you have 100+ questions and custom judge needs. Tools are means; the question bank is the asset. The asset is yours; tools are replaceable anytime.
The half-day schedule
| Time | Do | Produces |
|---|---|---|
| 9:00 – 10:30 | Mine 100 real inputs from production logs, run anonymization | Raw input pool |
| 10:30 – 12:00 | Hand-label pass conditions, first 50 (don't be greedy) | golden.jsonl v0.1 |
| 13:30 – 15:00 | Write graders.py + run.py, get them green | A working pipeline |
| 15:00 – 16:00 | Run the baseline, commit to git, wire up CI | baseline.json + green Actions |
| 16:00 – 17:00 | Write judge rubric v1 (only if open-ended questions), calibrate on human labels | judge_prompt.txt |
Pre-launch checklist
Check every box before release — miss one and don't claim you "have evals":
- Questions come from real logs, not invented; each has category and difficulty labels
- Rule assertions cover 60%+ of questions; a run finishes in under 10 minutes at acceptable cost
- Judge (if any) calibrated on 30+ human-labeled samples, agreement ≥ 85%, model version pinned
- Baseline committed to git; gate line (baseline − tolerance band) written down
- CI auto-runs evals on prompt/model changes; pipeline goes red on regression
- Production model uses an exact version number, not a latest-style floating alias
- Rollback switch genuinely works — flipped for real once, not "should work"
- Weekly 30-minute spot checks on the calendar; the failure → question pipeline exercised once
The first month's maintenance rhythm
Building it is just the start. Evals are a living thing — raise them on this rhythm for month one:
| When | Do | Goal |
|---|---|---|
| Week 1 | Read the report daily; learn the normal jitter range | Know "what counts as normal," don't flinch at noise |
| Week 2 | First spot check; turn failures into questions | Bank grows from 50 to 60+ questions |
| Week 3 | Deliberately make one "bad change" (e.g. delete a key instruction) and see if CI blocks it | Prove the gate is real, not decoration |
| Week 4 | Retro: which categories score lowest? Which to prioritize next month? | Build the "measure → add → re-measure" habit |
That week-3 sabotage drill — seriously, do it for real. A gate that has never blocked a bad change is one you won't truly trust; an untrusted gate is no gate. The drill costs half an hour and buys a year of peace of mind.
Closing: evals don't guarantee wins — they guarantee you know what you're betting
Back to the opening verdict. Evals won't stop your AI feature from ever breaking — the bank will always have blind spots, judges will always be biased, models will always change where you can't see. What they give you is exactly three things:
- Knowing what you're betting: every model swap and prompt edit comes with a number saying "this changed by how much."
- Knowing the moment you lose: not when users complain, but the instant CI turns red.
- Knowing how to win it back: versions pinned, prompts in git, the switch rehearsed — rollback is a one-minute job.
What a solo team can't afford to lose isn't money — it's trust. Users' trust in an AI feature, lost once, takes ten "seems fine"s to earn back — and it never really comes back. Half a day, one minimal eval setup. From now on, every model swap, you're holding cards — not hope.
Comments (0)
Related articles

Frank Michael Smith built the sports quiz game GeoSports in 12 hours with Claude Code: 79 players on day one, a million within a month, 15M+ plays and ~$47K ARR five months on. Why sports trivia won — plus five replicable judgments for indie developers, and three honest caveats.

Stars don't pay the bills. This guide lays out four proven monetization routes for indie open-source developers — donations, open core, dual licensing, and paid hosting — with a decision table, pricing funnel, and the hard lessons on drawing the free/paid line early.

Taxes aren't an "after you scale" thing — they're an "after your first dollar of revenue" thing. A practical field guide for indie developers: three real-world cautionary tales, US/EU/China compliance playbooks, the minimum one-person toolchain, and a 30-minute monthly reconciliation SOP.