After the AI Writes the Code: A Practical Code Review Workflow for Vibe Projects
The faster AI writes code, the more review matters. Four layers: diffs for logic (boundaries, errors, concurrency — plus auth, payments, SQL, encryption, secrets), runtime for behavior (type checks, lint, security scans go green first), AI for first-pass screening (a second model reviews, humans read only flagged parts), humans for the final call (AI never clicks merge). Includes commit norms, PR template, branch protection, rollback plans.

Everyone who codes with AI knows this moment: the agent spits out 500 lines, the tests run, everything is green. You exhale, click merge. Three days later production blows up — an unhandled edge case here, a secret written into the logs there.
The problem is not that AI writes too slowly; it is that you review too casually. Vibe coding drove the cost of writing code to zero, but the cost of reviewing it did not follow — it went up: five times the code volume, laced with logic you have never read a line of. This guide gives you a battle-tested workflow with one core idea: layered review — diffs for logic, runtime for behavior, AI for first-pass screening, humans for the final call. Every layer comes with a concrete action checklist. Just follow it.
Why review at all: the agent's three classic failure modes
First, align on the premise: reviewing is not about distrusting AI; it is about the fact that AI fails differently from humans. Humans make careless mistakes; AI makes confident mistakes — code that reads beautifully while hiding landmines. Three patterns show up most.
First: logic that looks right but is wrong. The hardest to catch. The code runs, the tests pass, but an edge case is wrong. A pagination offset written as offset = page * size when it should be (page - 1) * size; timezones mixed, UTC and local time used interchangeably; two concurrent requests decrementing inventory with no lock. AI is exceptionally good at writing code that looks right — clean variable names, thorough comments, tidy structure — and your eyes glide right over it. Review such code counterintuitively: the prettier it looks, the more carefully you read each branch condition.
Second: hidden security holes. To get things "working," AI takes shortcuts: API keys in source code, SQL built with string concatenation, user input interpolated straight into shell commands, an auth middleware "to be added later" that never gets added. More insidious is dependency poisoning — AI may pull in an npm package you have never heard of with malicious code inside. Eyeballing diffs cannot cover all of this, which is why we run automated scans before human eyes get involved.
Third: over-engineering. You ask for "a user feedback form" and get a plugin-based, extensible form engine supporting 10 backends with three layers of abstraction. No bugs, but triple the maintenance cost. Be willing to delete in review: ask "will this abstraction actually be used in the next three months?" If you cannot answer yes, have it simplified.
Keep these three in mind — every checklist below is designed around them.
Make code reviewable first: small commits, one PR per concern
The first step of review is not reading code; it is making code worth reading. An agent's default working style is a reviewer's nightmare: 2,000 lines in one shot, mixing a feature, a refactor, and a bug fix, with a commit message that says "update." No one can review that diff well, so set the rules before the agent starts.
Three iron rules — put them directly in your prompt or AGENTS.md:
- Small commits. Commit after each independent unit of work; one commit does one thing. A PR's diff should ideally stay under 400 lines — beyond that, human review quality falls off a cliff.
- One PR, one concern. Features, refactors, and bug fixes go in separate PRs. A mixed PR gets sent back for the agent to split. No exceptions, no softening.
- Change notes included. Every PR must say what changed, why, and how it was verified. Have the agent write this itself — if it cannot explain the change clearly, it did not think it through.
A reusable instruction template for your agent: "After finishing the work, split it into small commits by feature, one concern per commit; open a separate PR per feature with the what, why, and verification method in the description; keep each PR's diff under 400 lines — split if it exceeds that." Lock this in and your later review workload drops by at least half.
One more habit: ask the agent to flag "things I am unsure about" in the PR. AI usually knows where it is guessing — an API behavior written from memory, an edge case it is not sure about. Make it mark those spots explicitly, and review can go straight to them. Maximum efficiency.
The diff-level checklist: sweep everything first, then camp on the danger zones
Diff review happens in two passes. First pass: read through, catch logic issues. Second pass: danger zones only, catch security issues. Do not mix them — human attention does one thing at a time.
The general checklist for the first pass, item by item:
- Boundary conditions: empty arrays, null, zero, negatives, overlong strings, first and last pages — AI's favorite places to crash.
- Error handling: does every external call (API, database, file) have catch/finally? Could error messages leak sensitive data into logs? Any empty catch blocks swallowing exceptions?
- Concurrency: any races on shared state? What happens when two requests write the same row? Do you need locks, transactions, or idempotency keys?
- Resource cleanup: file handles, DB connections, timers, event listeners — anything leaking? Pay special attention to long-lived objects AI creates.
- Naming and comments: do comments explain the why, not the what? Do comments match the code? AI comments often describe the logic it thought it wrote rather than the actual logic — when they disagree, trust the code and fix the comment.
The second pass covers five danger zones where AI's error rate is highest — worth reading line by line:
- Auth and authorization: is every route protected? Any "allow for now, TODO later"? Is the permission check done server-side (client-side checks count as nothing)?
- Payments: are amounts computed in the smallest unit (cents)? Is there idempotency against double charges? Does the backend decide the price, or does it trust whatever the frontend sends?
- SQL: is everything parameterized? Search for traces of string-concatenated SQL — zero tolerance.
- Encryption: any home-grown crypto? How are keys and IVs managed? Are passwords hashed (bcrypt/argon2) rather than reversibly encrypted?
- Secrets and config: any hard-coded keys, tokens, or passwords in the code? Did config files get committed to the repo?
Practical trick: keep search open during review and grep the whole repo for TODO, FIXME, sk-, api_key, and password. AI's TODOs often hide unfinished work, and one search takes 30 seconds.
Runtime verification: let machines look first
Before human eyes touch the diff, let machines run everything they can. Order matters: automated gates go green first, humans look after; if it is not green, humans do not look.
- Type check + lint, all green. Run
tsc --noEmitfor TypeScript projects,mypyor at leastrufffor Python. AI-written code loves the "any" escape hatch — type checking flushes out many hidden null problems. - Tests, all green. Have the agent write tests for new features, and make sure they actually ran. Check that tests assert behavior rather than implementation details, and that edge cases are covered.
- Dependency security scan. Run
npm auditorpip audit; upgrade or replace on high-severity findings. For unfamiliar dependencies AI introduced, glance at weekly downloads and the last release date — a package published three days ago with double-digit downloads gets replaced, no debate. - Secret scanning. Scan with gitleaks or similar, and manually search prefixes like
-----BEGIN(private keys),sk-, andxoxb-. Once a secret enters git history, rotating it matters more than deleting it. - Run it and click through. The last step is always manual: run the feature and click through it twice — once as a normal user, once as a troublemaker. AI's happy path is usually fine; the devil lives in the error paths.
Turn these five steps into CI that runs on every PR. The classic vibe-project mistake is "CI later" — which means CI never. Spend one afternoon setting it up and save half an hour on every future PR.
AI reviewing AI: a second model does the first pass
When diffs get large, reading every line is neither realistic nor necessary. That is when a second AI does the first-pass screening — and note the word second: preferably a model from a different vendor with a different architecture. Having the code-writing model review its own code is like letting the student grade their own exam; the blind spots overlap.
The review agent's prompt must be specific — "take a look" does not cut it. A usable template:
You are a code reviewer. Review this PR diff and report three categories of issues: 1) logic bugs — concrete problems with boundary conditions, concurrency, or error handling, with file and line numbers; 2) security issues — authentication, SQL injection, hard-coded secrets, sensitive data leaks; 3) over-engineering — abstractions and dependencies that could be deleted. Sort each category by severity, report only issues you are confident about, do not speculate. Mark anything uncertain as "needs human confirmation."
The flow: the review agent produces its report; you read only what it flagged red (high severity), plus your own second-pass danger zones from the diff checklist. Roughly 80% of an AI report is correct, 15% is false positives, 5% it missed entirely — your job is that 20%. In practice, for a 400-line PR, the part needing careful human reading is usually under 50 lines.
But there is a red line: AI never makes the final merge decision. The review agent can flag issues and suggest fixes, but the hand that clicks merge must be human. This is not sentimentality; it is accountability — when something breaks, you cannot tell users "the AI told me to merge it." Write this into your team norms and into your own habits.
Merge gates: PR template, main protection, rollback plan
Review workflows need mechanical guardrails, not willpower. Three gates, all required.
Gate one: PR template. Drop a .github/pull_request_template.md in the repo forcing every PR to answer four questions: what changed, why, how it was verified, and what is uncertain. The agent picks it up automatically when opening a PR, and empty sections are visible at a glance.
Gate two: main branch protection. Turn on branch protection: CI must be green to merge, at least one human approval required, no direct pushes to main. Three clicks in GitHub or GitLab, and they stop 90% of "oops merges." Agents with push access are the norm in vibe projects — branch protection is the last physical barrier.
Gate three: rollback plan. Before merging, answer: if this breaks, how do we roll back in five minutes? git revert is the baseline; database migrations must be reversible (every up has a down); wrap big features in feature flags so you can kill the switch before fixing the code. Have the agent write one line about rollback in the PR description — a PR that cannot state its rollback plan is too big and needs splitting first.
Make it a habit: revisit, record, feed back
The final link is compounding. Spend 20 minutes a week on three things, and your agent will make half as many mistakes in three months.
First, review old code. Pick a PR merged two weeks ago and re-read its diff. Ask: what did I miss back then? Many issues are invisible in the moment and obvious in hindsight. Write down what you missed — that is the upgrade source for your checklist.
Second, record the agent's mistake patterns. AI mistakes follow patterns, and they are remarkably stable: one agent always forgets timezones, another always swallows exceptions, a third concatenates SQL strings every time. Keep a document (call it AGENTS.md or review-notes.md) and log each recurring pattern with an example.
Third, feed patterns back into prompt specs. Recording is not the goal; prevention is. Translate high-frequency patterns into prompt rules in the agent's system instructions: "store all timestamps in UTC, convert to local time only for display," "no string-concatenated SQL, always parameterize," "empty catch blocks are forbidden." Next time the agent writes code, it steers around those pitfalls on its own — the checklist you tuned by hand becomes its writing spec.
Once this loop spins up, you get a virtuous cycle: the more carefully you review, the more patterns you record; the more patterns feed back into specs, the cleaner the agent's code; the cleaner the code, the faster the review. The end state of vibe coding is not "AI writes, humans don't look" — it is "AI writes, humans set the standard, AI writes to the standard." The review workflow is how you set the standard.
Related articles

Prompts in vibe projects live in code strings, admin text boxes, and docs — changed live, version unknown when things break. This guide shows how to treat prompts like code: a prompts/ layout, YAML frontmatter, semantic versioning, PR reviews, canary rollouts with one-click rollback, plus an evals baseline — and a real war story: one added sentence cost 12 points of classification accuracy.

Getting signups is only the start — users churn by day 3 and you have no horn to call them back. This guide covers notification systems for vibe projects: channel selection, email with Resend from day one, SPF/DKIM/DMARC done right, when SMS is worth the money, frequency caps and unsubscribe, retries and dead letters, plus a launch acceptance checklist.

Cloudflare's Birthday Week blog makes the case plainly: GitHub was designed for humans writing code; the agent era needs the collaboration layer reinvented. Artifacts enters open beta with a repo for every agent, plus a developer competition — $25,000 in credits for first place, deadline October 14. This is the first time a major infra vendor has put 'infrastructure for agents writing code' on the table as a public proposition.