Testing AI-Generated Code: From "It Runs" to "Ship It"
The payment webhook that demoed perfectly and exploded three days after launch: four correct postures for having AI write tests, three high-voltage lines requiring human review (money/permissions/deletion, concurrency, migrations), and a minimal pipeline for tonight.

With AI writing code, "it runs" is easy. You describe the requirement, it spits out code, you click around, it works — dopamine hit. But "ship it" is a different matter: at 3 AM, woken by alert texts, you discover that the most dangerous thing about AI-generated code is precisely where it "looks fine." It never tires, and it never doubts itself; the bugs it writes are equally confident.
This piece sells no anxiety, only method: first a true story of a "perfect demo, exploded three days after launch" incident, then the correct posture for having AI write tests (wrong posture means the more tests you write, the more falsely reassured you are), then the three high-voltage lines where tests must be human-reviewed, and finally a minimal test pipeline you can set up tonight. The goal: trade the luck of "it runs" for the certainty of "ship it."
First, the Incident: Perfect in Demo, Exploded on Day Three
Lao Zhou (a pseudonym), an indie developer, used AI to build a paid-knowledge site. The core flow: user pays → Stripe webhook → access granted → email sent. Before launch he personally clicked through the full flow 5 times: pay, receive, unlock, get the email — all green. Demoing to friends, everyone said "solid." On day three after launch, at 2 AM, he was jolted awake by a flood of texts: 20+ users had received duplicate access emails, and 3 were double-charged.
The postmortem is a textbook AI-code incident: Stripe webhooks have a retry mechanism — if the first callback is slow, Stripe sends it again, and again. The AI-generated webhook handler had no idempotency check: one request in, one course unlocked, one email sent, one ledger entry. Zhou's 5 manual clicks were all "clean" single requests — a retry storm could never be caught that way. And when writing the code, the AI had written zero tests for the failure path — every test case was "one happy payment, assert access granted." Happy path: all green.
The incident cost over 4,000 yuan in refunds and apology red packets. Zhou's summary cut deep: "I had the AI write the code and then had the AI write the tests — that's letting the student write the exam and grade it. A perfect score is guaranteed; passing is accidental." From that day he set an iron rule: AI may write tests, but "what the tests are actually testing" is decided by a human.
Why is AI especially prone to missing failure paths? Three reasons: first, survivorship bias in training data. Tests in open-source repos are already mostly happy path, so the "tests look like this" that AI learned is mostly happy path too. Second, AI has never been woken at 3 AM. Human programmers write idempotency and retries because they've been burned; AI has never been burned — it doesn't know pain. Third, your prompt never asked. "Write a payment callback" is a feature description, not a quality description — if you don't mention failure paths, the AI assumes you don't care. Failure-path tests are therefore always something you (not the AI) must proactively demand.
The Correct Posture for Having AI Write Tests
First, correct a misconception: having AI write tests isn't wrong — the posture is. Four correct postures:
- Posture one: tests are the spec — write assertions before implementation. Don't wait until the code is done and ask the AI to "add tests" — by then it will just write assertions that conveniently pass against the existing implementation. The right order: you write acceptance criteria in plain language first ("a duplicate callback for the same order must grant access only once"), have the AI translate them into test cases, run them, watch them go red, then have the AI write the implementation to turn them green. Here the test's role isn't "verifying code" — it's "locking the spec." Once the spec is locked by tests, the AI's later refactors can't improvise wildly.
- Posture two: give the AI "counterexample seeds" and let it proliferate. AI is bad at "inventing tricky edge cases from thin air" but excellent at "proliferating from examples." You list 5 counterexamples you can think of (empty strings, negative numbers, oversized input, concurrent duplicate requests, timezone boundaries), then instruct: "generate tests for each counterexample, then proliferate 10 more of the same kind." In practice, about 30% of the AI-proliferated edge cases are ones you never thought of — and that 30% is usually where production incidents come from.
- Posture three: property-based tests instead of "example" tests. Traditional tests are "examples": input A, expect output B. AI-written example tests have a fatal tendency — the examples always pick the smoothest path. Property-based testing flips it: define rules that hold "for any input" (e.g., "serialize-then-deserialize must equal the original," "sorted output must be ordered") and let the framework hurl thousands of random inputs at them. For AI-generated parsing, transformation, and computation code, property tests catch far more bugs than example tests. Having the AI write the properties and generators plays exactly to its "tireless" strength.
- Posture four: coverage is a reference, not a goal. 90% line coverage sounds beautiful, but it's trivially easy for AI-generated code — the AI happily generates filler tests that "call it and assert true." What actually deserves scrutiny is the untested failure branches in branch coverage: catch blocks, retry logic, fallback paths, idempotency checks — precisely the high-incident zones in production. Require the AI to list "uncovered exception branches" separately in its test report, and you decide each one: add a test, or confirm "this branch genuinely can't happen."
- Posture five: mutation testing — have the AI sabotage itself. After code generation, open a new task: "deliberately break this code in 10 places (wrong conditions, deleted null checks, flipped comparisons) and see whether the existing tests catch them." The "breaks" they miss are your blind spots. Mutation testing tests the tests — ideal for AI-generated code, because AI's tendency to write tests "along the implementation" makes blind spots systematic. Run monthly; blind spots shrink.
Beyond postures, a mindset: treat "test went red" as good news. Many panic when AI-generated tests go red, thinking "the AI wrote it wrong." Wrong — a red test is catching an implementation problem early: red on your dev machine beats red in production. Building the "red-then-green" muscle memory is the vibe coder's rite of passage from "toy" to "engineering."
Which Tests Must Be Human-Reviewed: Three High-Voltage Lines
AI writes tests, humans review tests — but human attention is finite, so review firepower must concentrate on the most dangerous ground. Three high-voltage lines; touch any one and the test cases get a line-by-line human review:
- High-voltage line one: money, permissions, data deletion. Any code touching fund flows (payments, refunds, billing), permission changes (elevation, sharing, going public), or data deletion (soft delete counts) gets every assertion human-reviewed. Review what? "What is this assertion actually testing": does assertEqual(balance, 100) test "the right money was deducted," or just "the code ran"? AI loves "self-congratulatory assertions" — e.g., calling the API, grabbing the result, then asserting the result equals what it just grabbed (assertEqual(result, result) in its many disguises). Such tests are forever green and utterly worthless. For high-voltage code, a human must confirm "every assertion maps to a real business rule."
- High-voltage line two: concurrency and race conditions. AI is strong at single-threaded logic and terrible at concurrency — its training data simply contains few samples of "correctly handled concurrency." Any scenario involving "simultaneous," "duplicate," or "timeout retry" (flash-sale inventory deduction, webhook idempotency, distributed locks) must be human-checked for whether the tests genuinely construct concurrency: multiple threads/processes hammering it? A simulated retry storm? If someone had glanced at Zhou's test list and seen the five characters "no concurrency tests," the tragedy would never have happened.
- High-voltage line three: data migration scripts. Migration scripts run once and can't be un-run if wrong. For AI-generated migrations, tests must include: a dry-run mode (prints SQL, executes nothing), a rollback script, and a record of "ran once against a copy of a production snapshot." No negotiation here — an unreviewed migration script never touches the production database. Iron rule.
Beyond the high-voltage lines, one universal review heuristic: check whether the AI's tests contain any counterexamples at all. A test file with only happy paths is untested. Quick standard: if "exception," "boundary," and "concurrency" cases together make up less than 30% of the test file, send it back for rewrite.
When reviewing assertions, remember the "three questions": one: if this assertion fails, which business rule can I pinpoint as broken? A failure message of "expected true but was false" says nothing — good assertions carry business semantics. Two: if I delete the implementation, does this test go red? Comment out the code under test and re-run; tests that stay green can be deleted outright. Three: does this test have a "counterexample sibling"? Next to every happy-path assertion, is there a matching exception or boundary assertion? If not, it's running naked. Can't answer all three — send it back.
One more cold fact: AI-written tests usually pass at a suspiciously high rate on first run — near 100%. Don't be fooled: it proves the tests and the implementation came from the "same brain" with fully overlapping blind spots. The healthy signal is 20-30% red on first run, then green after fixing the implementation. An all-green suite equals untested. Next time you see all green, suspect first, celebrate later — all-green is the color that most deserves suspicion.
A Minimal Test Pipeline You Can Set Up Tonight
Theory is easy; landing it is hard. This pipeline is "minimal viable" — a solo project can set it up tonight:
- Step one: set the rule for the AI — "no tests, no merge." Add one line to your project's AGENTS.md (or equivalent agent instruction file): every AI-generated PR must ship with test cases, and tests must go red-then-green (assertions first, watch red, then implementation, watch green). Written into the agent's "factory settings," it's a hundred times more reliable than asking verbally each time.
- Step two: pre-commit runs unit tests locally — must finish in 2 minutes. Use a git hook to auto-run unit tests before commit. The key constraint is "fast": over 2 minutes and humans will find ways around it. Slow integration tests go to CI; local runs unit tests only. AI-generated code must pass this gate before entering the repo.
- Step three: staging runs on "shadow traffic." Copy real production requests (sanitized) into staging and let the AI-generated code run under shadow traffic for 24 hours. Zhou's webhook incident would have surfaced in 10 minutes under shadow traffic — the real world's retry storms can never be simulated locally.
- Step four: monthly, have the AI red-team its own code. Open a fresh session with a full "attacker" persona: "You are a security researcher; this code is your target; find every exploitable vulnerability and write PoC tests." AI is surprisingly effective at finding bugs in its own (or another AI's) code — it's unburdened by the psychological defense of "my code is fine." Once a month, half an hour, frequently surprising (or alarming).
- Step five: keep a "test smells" list on the wall. Have the AI compile one: sleep(1000)-style hard waits, assertTrue(true), shared global state between tests, random numbers without seeds, twenty assertions in one test method… Every test review starts with a pass over the smells list. Smells aren't bugs, but they're where bugs breed — and AI-generated code is especially smell-rich, because it all "looks" correct.
My take: testing isn't distrusting AI — it's trading luck for certainty.
Some see testing AI-generated code as "not trusting AI." The opposite is true — it's precisely because we trust AI's output velocity that we need tests guarding the quality floor. AI development without tests bets "demo-day luck" against "post-launch certainty," and the odds of winning that bet decay exponentially as the codebase grows.
Deeper still: testing is the cleanest division-of-labor interface between human and AI. AI does the "tireless generating" — code, test cases, boundary proliferation; humans do the "non-outsourceable judging" — what this assertion is testing, whether this high-voltage line passes, whether this failure branch is acceptable. After his incident, Zhou said: "I used to see review as a burden. Now I understand: reviewing assertions is the last — and most important — craft of programmers in the vibe coding era."
One last counterintuitive note: the teams that write the most tests are usually not the slowest — they're the bravest at refactoring. Refactoring will be extremely frequent in the AI era (model upgrades, prompt tweaks, architectures torn down and rebuilt), and without tests as a safety net you won't dare let the AI touch old code — so it rots there. Tests aren't cost; they're the confidence behind every future "have the AI overhaul this."
So stop asking "should AI-written code be tested" — yes. But rephrase the question and the answer becomes more valuable: "In my test suite, how many assertions have I personally confirmed are testing the right thing?" That number is exactly how "shippable" your code is. Start tonight by adding three counterexample tests to your core flow.
Related articles

Every vibe project has the same darkly comic moment: your site goes white-screen and a friend tells you before your monitoring does. This guide builds a one-person-team error monitoring system: a 5-minute Sentry loop, error boundaries, report context design, backend structured logging, AI-call-specific protection, alert tiers, and a launch checklist.

On October 3, engineer Kevin Liao published a polemic that hit the HN front page: agent memory plugins are a lottery over RAG snippets; what agents need is a documentation workspace. The essay's diagnosis, its open-source Operator Memory plugin, the two strongest objections, and the minimal practice you can start tonight.

Gergely Orosz visited OpenAI, Anthropic, Cursor, and Ramp and wrote up the 2026 state of the industry: near-100% AI-generated code, agent PRs up ~10x in eight months, code review degrading into theater, the IDE declared legacy. Key takeaways plus three verdicts and four actions for vibe coders.