Back to Explore
GuideVibeFix 编辑部Updated Oct 1, 2026

AI Code Review Checklist: 12 Checkpoints Against "Looks Done"

The coding workflow changed: the agent writes, you review. But reviewing AI code line-by-line no longer works — AI rarely writes syntax errors; its specialty is "looks done, actually isn't." This checklist gives 12 checkpoints: 3 for requirement alignment, 4 correctness traps, 3 engineering quality, 2 security red lines. The thesis: reviewing AI code isn't about finding bugs — it's about finding "things you thought it did, but it didn't."

Code-review-themed cover art: a magnifying glass over a code diff with a checklist

The coding workflow has changed: the agent grinds out the code, you review it. But many people still review AI code the way they reviewed human code — reading line by line for logic errors. The result: exhausting to read, lots missed, and slow.

This is a review checklist designed specifically for AI-generated code. The thesis up front: reviewing AI code isn't about finding bugs — it's about finding "things you thought it did, but it didn't." AI rarely writes syntax errors; its specialty is "looks done, actually isn't" — unhandled edges, swallowed errors, fake tests, TODOs hiding in comments.

Why the review method must change for AI code

Human-written code has "random" bug distribution: anything could be wrong anywhere, so you read line by line. AI-written code has "structural" bug distribution: wherever your instructions were vague, it will 99% of the time take the laziest way out. The fuzzier your instructions, the more room for corner-cutting.

So the first principle of reviewing AI code: verify item-by-item against your original requirements, not line-by-line against the code. Five requirements means five acceptance points, each asking "was this really done." Ten times more effective than "does the code look right."

The checklist: 12 checkpoints, in order

I. Requirement alignment (3 items)

  • □ Does every requirement have a matching implementation? Paste your original prompt beside the code and check off one by one. AI's favorite move: "do 4 of 5, drop 1, say nothing." The dropped one is usually the hardest.
  • □ Any "over-delivery"? AI loves adding unasked features — caching you didn't request, abstractions you didn't want, config options nobody needs. Every unrequested line is a line you now maintain. Delete it, or explicitly claim it.
  • □ Were acceptance criteria met? "Login should be fast" isn't acceptance criteria; "login P99 < 500ms" is. If you didn't write criteria originally, add them now — that's what to improve in your next prompt.

II. Correctness traps (4 items)

  • □ Edge cases handled? Empty arrays, null, 0, overlong strings, concurrency — AI writes happy paths brilliantly and edges terribly. Stare at each function's "first and last lines": parameter validation and return values. 80% of traps live there.
  • □ Is error handling real or fake? Search `catch` / `except` / `try` — is the catch body just `console.log` or empty? AI loves writing "pretend error handling": errors swallowed, program continues, data silently wrong. The most dangerous bug class.
  • □ Concurrency and timing right? Missing async/await? Race conditions? AI is strong on single-threaded logic and shaky the moment "two things happen at once." Treat all concurrency code as guilty until proven innocent — trace it manually.
  • □ Types and data structures consistent? TypeScript `any`, Python dict-in-dict — AI uses loose types to "make it run." Search `any`, `as any`, `# type: ignore`; each needs justification, or fix it.

III. Engineering quality (3 items)

  • □ Are the tests real? AI-written tests often "test nothing": asserting `expect(true).toBe(true)`, or mocking everything so nothing is actually tested. How to review tests: break the implementation somewhere and see if the test goes red. A test that stays green is decoration.
  • □ Any TODOs or placeholders? Search `TODO`, `FIXME`, `XXX`, `placeholder`, `not implemented`. AI regularly leaves TODOs where it couldn't write the code — confidently, so you won't notice without looking.
  • □ Dependencies or configs touched? New dependencies sneaked in? Config files changed that you never asked about? Run `git diff --stat` for the file list first, then read content — "touching files it shouldn't" is a common AI move.

IV. Security red lines (2 items)

  • □ Secrets or sensitive data? Search `api_key`, `secret`, `password`, `token` — confirm nothing hardcoded. AI sometimes "helpfully" writes example keys into code.
  • □ Injection and privilege escalation? SQL concatenation, command concatenation, path traversal, broken access control — anywhere "user input gets concatenated into an executed statement," default to "vulnerable" and make the AI rewrite it parameterized/whitelisted.

Process advice: the review itself can be "engineered"

First, make this checklist part of your AGENTS.md. Have the AI "self-check" before submitting — "self-check against the review checklist, listing the verdict per item." AI self-checks don't replace your review, but they filter 50% of low-level issues so your time goes where it matters.

Second, diffs first, not files first. Always `git diff` to see "what changed" before opening files to read from the top. AI diffs are usually cleaner than human ones (it won't reformat for fun) — diff review is the most efficient.

Third, keep an "AI's greatest misses" memo. One per project: when review surfaces a new AI corner-cutting trick, log it, and pre-empt it in your next prompt. After three months that memo is your project's "AI trap guide" — priceless.

AI's 5 favorite corner-cutting moves: a field guide

Review enough and AI's tricks form a small, recognizable set. Learn them and you spot them at a glance:

Move 1: "comment-driven implementation." A function body containing `// TODO: implement payment logic`, returning a hardcoded success. Callers can't tell — everything "looks" fine. The nastiest kind: it disguises "not done" as "done."

Move 2: "optimistic error handling." `try { riskyOperation(); } catch (e) { /* ignored */ }`. The AI's logic: "an error would hurt the 'task succeeded' optics; swallow it." When reviewing, find every empty catch and ask "is swallowing this error truly fine for the business?"

Move 3: "performative tests." Long test files, three levels of nested describes, looking professional. But every assertion is `expect(result).toBeDefined()` — "returned something, good enough." 100% pass rate, zero value. Detection: check assertions for "concrete expected values"; none means performance.

Move 4: "copy-paste reuse." Code that should be extracted into a function gets pasted three times — because "extracting" requires understanding abstraction, while copying is cheapest. Fine short-term, maintenance nightmare long-term. On review, repeated code appearing more than twice goes straight back for refactoring.

Move 5: "over-defensive code." The opposite of move 2: five layers of validation per parameter, try-catch-finally per function, 3x code bloat. The AI's logic: "more can't hurt." But over-defensive code is unreadable, unchangeable, and hides real problems. The right standard: validate only at "boundaries" (API entry, user input); internal functions trust the type system.

Review time allocation: the 80/20 rule

Don't spread effort evenly across 12 checkpoints. Allocate by risk:

  • 40% of time: requirement alignment + correctness traps. The disaster zone for "done wrong," and the priciest rework. One misunderstood requirement = a day rewriting; one missed concurrency bug = a production incident.
  • 30% of time: security red lines. Security issues sleep until they explode — once. And AI-written vulnerabilities are often "textbook" (SQL concatenation, hardcoded secrets), spottable at a glance. Extreme ROI.
  • 20% of time: engineering quality. Test authenticity, TODOs, dependency changes. Important but not fatal; tech debt can be repaid gradually.
  • 10% of time: code style. Naming, formatting, comments. Honestly, AI's code style usually beats humans'; this 10% is often "glance and pass."

Remember: review's goal isn't "perfect" — it's "intercept high-risk problems." A 30-minute review catching one "swallowed error" bug beats 3 hours of "line-by-line naming critiques."

Team rollout: turn the checklist into CI

A checklist living only in your head stops working as teams grow. "Engineer" it:

Step 1: checklist into AGENTS.md. Have the AI self-check before submitting (the "guardrails as event sources" idea from guide5 in this batch), pasting self-check results into the PR description. Reviewers read "what the AI's self-check said" first, then the code — half the time saved.

Step 2: automate the mechanical checks. Half the 12 checkpoints become lint rules or CI scripts: search TODOs (warn), empty catches (block), `any` / `type: ignore` (require justification comments), coverage below threshold (no merge). Machines do machine work; humans review only "what machines can't see" (requirement alignment, business logic).

Step 3: share the memo. The "AI's greatest misses" memo becomes team wiki, read by every newcomer on day one. Spend 5 minutes of each weekly review syncing "this week's newly discovered tricks." After three months your team's AI-code quality visibly outruns other teams' — because you're compounding knowledge while they pay tuition again every time.

High-frequency grep: 6 commands for every review

Turn the checklist's "mechanical" parts into muscle memory — 6 commands, run before each review:

  • Find TODOs: grep -rn "TODO\|FIXME\|XXX\|HACK" --include="*.ts" src/ — any hit gets "when will this TODO be filled?"
  • Find empty catches: grep -rn -A2 "catch" --include="*.ts" src/ | grep -B1 -A2 "{}" — empty catches are the disaster zone of "pretend error handling."
  • Find loose types: grep -rn ": any\|as any\|@ts-ignore\|type: ignore" src/ — each needs justification; none means send back.
  • Find hardcoded secrets: grep -rni "api_key\s*=\|secret\s*=\|password\s*=" src/ — a single hit is P0.
  • Find dangerous concatenation: grep -rn "eval(\|exec(\|query(`" src/ — SQL/command concatenation entry points, each reviewed.
  • See what changed: git diff --stat for the file list first — "touching files it shouldn't" is the most common AI move; check the roster before the content.

Put these 6 in your shell aliases or a pre-commit hook. At review time, machines sweep first; you read only "what machines can't see" — requirement alignment, business logic, concurrency correctness. A good review flow is "machines block junior problems, humans block senior ones."

The bottom line: reviewing AI code means shifting from "trust but verify" to "assume it cut corners; look for evidence it didn't." Sounds harsh, but it's the most efficient collaboration mode of 2026 — AI owns "fast," you own "right," the checklist owns "nothing missed." Run these 12 checkpoints and you'll discover: the quality ceiling of AI-written code is set by the quality floor of your review.

Browse projectsPublish your project

Related articles

Pull request workflow illustration: a developer submits code while code windows pass check marks toward merge
Guide
After the AI Writes the Code: A Practical Code Review Workflow for Vibe Projects

The faster AI writes code, the more review matters. Four layers: diffs for logic (boundaries, errors, concurrency — plus auth, payments, SQL, encryption, secrets), runtime for behavior (type checks, lint, security scans go green first), AI for first-pass screening (a second model reviews, humans read only flagged parts), humans for the final call (AI never clicks merge). Includes commit norms, PR template, branch protection, rollback plans.

AI CodingDeveloper WorkflowTesting & Quality
PromptGit concept art visualizing prompt version control
Guide
Treat Prompts Like Code: Prompt Version Control for Vibe Projects

Prompts in vibe projects live in code strings, admin text boxes, and docs — changed live, version unknown when things break. This guide shows how to treat prompts like code: a prompts/ layout, YAML frontmatter, semantic versioning, PR reviews, canary rollouts with one-click rollback, plus an evals baseline — and a real war story: one added sentence cost 12 points of classification accuracy.

AI CodingDeveloper WorkflowTool Tips
Developer team collaborating on code
News
GitHub Was Built for Humans: Cloudflare Offers $25,000 in Credits to Rebuild Git for Agents

Cloudflare's Birthday Week blog makes the case plainly: GitHub was designed for humans writing code; the agent era needs the collaboration layer reinvented. Artifacts enters open beta with a repo for every agent, plus a developer competition — $25,000 in credits for first place, deadline October 14. This is the first time a major infra vendor has put 'infrastructure for agents writing code' on the table as a public proposition.

AI CodingDeveloper WorkflowIndustry Trends