Adversarial Review

The plain-English guide to adversarial review

AI coding tools can now build real software for people who have never written a line of code. That is a genuine shift. It also creates a new problem: the same tool that wrote the code is usually the one telling you it works.

This guide is about fixing that one problem. It is written for people who build with AI but don’t review code for a living. No jargon you don’t need, and nothing here requires you to read a diff.

The one idea

The AI that wrote the code should never be the one that grades it.

That’s it. Everything else in this guide is how to put that idea into practice.

When you ask an AI to build something, it spends the whole session collecting reasons to believe its own work is correct. It wrote the plan, it wrote the code, it often wrote the tests, and it read its own success messages. By the end it is the least objective reviewer in the room. It isn’t lying to you. It is doing what any author does: reading what it meant to write instead of what it wrote.

An adversarial review flips that. A second reviewer starts from nothing, is told to assume the work is wrong, and is only allowed to report. It isn’t allowed to edit anything, it can’t explain away, and it can’t lean on the builder’s reasoning, because it never saw it.

Why AI-written code needs this more, not less

Human developers make mistakes too. But AI-written code fails in a few specific, repeatable ways that are worth knowing by name.

1. “Done” that isn’t done. The AI finishes most of the request, quietly leaves a piece for “later,” and reports success. The missing piece is often mentioned only in a code comment you’ll never read: a TODO, a stub, a feature built but switched off.

2. Tests that test nothing. When the same AI writes the code and the tests, the tests tend to confirm what the code already does. They pass. They just don’t prove the thing you care about.

3. Silent failures. Errors get caught and ignored so the app doesn’t crash. The app keeps running, and nobody ever learns that the email didn’t send or the record didn’t save.

4. The happy path only. It works perfectly for the first, clean, expected input. It breaks on the empty field, the duplicate click, the customer whose data was created before the change.

5. Confident summaries. The end-of-session summary is the most persuasive thing the AI writes, and the least checked.

None of these are exotic. They are the normal shape of shipping fast with AI. The fix isn’t to stop. It’s to add one step.

The three rules of a good adversarial review

Fresh eyes. The reviewer must not share the builder’s context. In Claude Code, that means a separate agent that starts clean, not the same chat asking itself “are you sure?” Asking the builder to double-check its own work is like asking someone to proofread a letter they just finished writing.

Evidence, not assertions. “Tests pass” is a claim. “Here is the line where a failed payment is ignored” is evidence. A good review points to the exact file and line, or the exact command and what it printed.

It can only report. The reviewer doesn’t touch the code. (Our skill enforces this: the review agent has no permission to edit files.) The moment the reviewer starts fixing things, it becomes a second builder with the same blind spots. Findings go back to the builder, the builder fixes, and the reviewer runs again.

How to run one in about ten minutes

You can do this today with Claude Code and our free review skill. You’ll need a recent version of Claude Code (2.1.218 or later); if you’re not sure, ask Claude Code to update itself.

  1. Download the skill from the free skill page. It’s a single folder with one file in it.
  2. Install it by placing that folder in your Claude Code skills folder: ~/.claude/skills/ for every project on your computer, or .claude/skills/ inside one project. (Anthropic’s documentation on skills explains both.)
  3. When the AI says it’s finished, type /adversarial-review followed by one sentence describing what you asked for. For example: /adversarial-review add a cancel button that stops the subscription and emails the customer.
  4. Read the verdict. You’ll get one of three verdicts: SHIP, FIX FIRST, or STOP.
  5. If it isn’t SHIP, copy the “paste this to the builder” section back into your main chat. When the builder says it’s fixed, run the review again.

The skill runs in its own separate agent, so it never sees the conversation that produced the code. That separation is the whole point.

What to look for, even if you can’t read code

You don’t need to read code to ask good questions. These five areas are where AI-written changes most often hurt real people. If a change touches any of them, review it every time.

Money. Anything that charges, refunds, pays out, or calculates an amount. Ask: can this run twice? What happens if the payment fails halfway?

Who can see what. Logins, permissions, teams, accounts. Ask: could one customer ever see or change another customer’s data?

Deleting or overwriting. Ask: can this be undone? Is it limited to exactly the records it should touch?

Half-finished work. Ask: is every part of what I asked for actually reachable by a real user today, or is some of it behind a setting, a flag, or a “later”?

Silent errors. Ask: if this fails, who finds out, and how?

A useful rule of thumb: done means a real user can reach it. Merged isn’t done. Passing tests aren’t done. Built but switched off isn’t done.

Reading the verdict

SHIP means the reviewer looked hard and found nothing serious. It doesn’t mean the code is perfect. It means the obvious ways it could hurt someone have been checked.

FIX FIRST means real problems that a customer would hit. Send them back, get them fixed, review again. This is the most common result, and that’s fine. It’s the review doing its job.

STOP means something could lose money, leak data, or destroy data. Don’t merge it. Don’t let the builder talk you out of it in the same breath. Fix it, then review from scratch.

When one reviewer isn’t enough

For anything involving money, private data, or deleting things, use two reviewers that don’t share blind spots. A second AI model from a different company is a practical option: different training, different habits, different misses. If both independently pass the change, you can be far more confident than if either one does alone.

What this is not

Adversarial review is a habit, not a guarantee. It won’t replace a professional security audit, and it can’t test things it can’t see, like how your app behaves with real traffic or a real payment provider. Treat it the way a good editor treats a manuscript: it catches most of what the author can’t see, and it makes the author better over time.

The short version

  1. The AI that wrote it doesn’t grade it.
  2. A fresh reviewer, told to assume it’s wrong, only allowed to report.
  3. Evidence over claims, every time.
  4. Money, access, deletes, half-finished work, silent errors: review these always.
  5. Done means a real user can reach it.

Get the free review skill