Why the review bar goes up, not down
I use AI agents in my daily development workflow — often several in parallel. The output arrives fast and reads plausibly. That combination is exactly why review discipline matters more now than it did before. Plausible-looking code that's subtly wrong is harder to catch than obviously messy code a junior would have flagged themselves.
This is the checklist I run on diffs that didn't come from a human head. It's ordered the way I actually review — intent first, details later.
What AI is genuinely good at
Fair is fair. Where the tools are strong, I stop over-reviewing and move on:
- Boilerplate and scaffolding. Config files, standard CRUD endpoints, typed wrappers — usually correct and idiomatic.
- Pattern translation. "Port this class component to hooks," "convert this REST call to the fetch wrapper we use" — mechanical transformations are a sweet spot.
- Happy-path tests. Tests for the obvious cases come out decent, and a decent test suite is a floor I can build on.
- Explaining unfamiliar code. Asking an agent to summarize a module before I review changes to it has become routine.
The failures cluster elsewhere. Knowing the difference is the whole skill.
The checklist
Correctness vs. intent
- Does this solve the problem that was asked, or a nearby one? Agents optimize for the prompt's literal reading — verify it matches the ticket, not just the sentence you typed.
- Did scope creep in? Watch for drive-by refactors, "improvements" to adjacent code, and new dependencies nobody requested.
- Are the boundaries respected? Server code in server files, no
'use client'where the design says server component.
Edge cases
This is where the drift is worst. Hunt these branches deliberately:
- Empty collections — does the happy path assume
items[0]exists? - Null and undefined — optional fields treated as present,
?.chains that swallow a needed failure. - Off-by-one in pagination, slicing, date ranges.
- Concurrency: stale closures, missing
await, races between a request and its cancellation. - Retries and timeouts — often silently missing.
A concrete example of the genre. This passes review if you only skim:
// Looks fine. Fails on an empty cart — and on a cart
// that changes between the two reads.
const total = items.reduce((sum, i) => sum + i.price * i.quantity, 0)
await saveTotal(userId, total)
The fix isn't clever — it's the branch nobody wrote a test for:
if (items.length === 0) {
await saveTotal(userId, 0)
return
}
Security basics
- Input from the client is validated — on the server. An agent that puts validation only in a form component has given you a UI nicety, not security.
- No secrets in the diff. API keys and tokens show up in AI code with alarming regularity when they appear in the context you pasted in.
- Injection surfaces: raw SQL, unsanitized HTML,
dangerouslySetInnerHTML, shell command construction. - Authz checks exist at the data layer, not just the route.
Error handling
- Failures propagate. Look for
catchblocks that log and continue, or worse, return a default that renders as an empty success. - User-facing errors are actionable; silent failures are worse than thrown ones.
- Error types survive the journey — a generic
Errorwhere the UI needs a status code is a design miss.
Over-abstraction
- Is there an interface, factory, or wrapper with exactly one implementation? Delete it until there's a second.
- Is a utility function used once and three parameters deep? Inline it.
- Did a two-file feature grow a
utils/, atypes/, and a barrel export? Agents love structure; structure without a second consumer is ceremony.
Naming
- Names say what things are, not what the prompt asked for.
handleData2andprocessItems2are artifacts of generation, not design. - Booleans read as predicates (
isLoading,hasAccess). - Types are named after domain concepts, not after the response shape of one API call.
Tests
- Tests assert behavior, not implementation details (spy counts on internal calls).
- At least one test covers the edge cases from the section above — if none exist, write the empty-state test yourself before merging.
- Tests fail for the right reason. Temporarily break the implementation and confirm the test catches it.
Process tips that make the checklist cheaper
Keep diffs small. A 900-line agent diff is a review you will not finish carefully. When an agent produces one, I ask it to re-emit the work in slices — or I check out the slices myself. Review quality scales inversely with diff size.
Run the gates before reading. Lint, typecheck, tests. If those are red, there's nothing to review yet. If they're green, the human pass has one job: the things gates can't see.
Use a second agent as a critic — as a second pass. Asking an agent to critique a diff catches a surprising amount of low-hanging fruit. But it shares blind spots with the agent that wrote the code. It's a lint extension, not a reviewer replacement. The final read is mine, because the merge is mine.
The honest summary
AI writes the first draft faster than I do, and I've made peace with that. What it hasn't earned — what I won't hand over — is the judgment about whether the draft should ship. The checklist above is how I keep that judgment cheap enough to run on every diff, every time.
Trust the output enough to use it. Verify it like you'll be the one explaining it in production. Because you will be.