Adversarial Review

The success message is not the evidence

On one software project, an AI agent was asked to send an announcement email to a list of users. It built the email, sent it, and reported back: sent successfully, no errors.

The main link in the email was broken. Clicking it went nowhere.

Nothing in the pipeline was lying. The email was, in fact, sent successfully. The agent accurately reported what its tools told it. The only problem was that every one of those signals answered a different question from the one that mattered. “Did the email send?” is not the same question as “Will a reader who clicks this get what the email promised?”

The only reason anyone found out was that a person opened the email and clicked.

A claim versus evidence

This is the core distinction in reviewing anything an AI builds:

  • A claim is something a tool or a person says happened. “Tests pass.” “Deployed.” “Sent.” “Fixed.”
  • Evidence is the thing itself, observed directly. The page loading. The email in an inbox, clicked. The row in the database with the right value in it.

AI agents are good at producing claims. They run a command, read the output, and summarize it. That summary is usually accurate about the command. It’s often silent about the outcome you actually care about.

Here are a few shapes this takes:

“Tests pass.” True, and the tests check that a function returns something, not that it returns the right thing.

“Deployed successfully.” True, and the new version is live on a preview address, not the one your customers use.

“Sent.” True, and it went to the test address, or the link is dead, or it landed in spam.

“Fixed.” True for the one record that was reported, and the same bug is still sitting in every other record created the same way.

In each case the success message is real. It’s just not the evidence.

Make the agent open the artifact

The fix is a habit, and you can write it straight into your instructions to the AI:

Never report that something works because a tool said it did. Open the thing.

In practice that means:

  • For a web page: load the real address and check the content that changed.
  • For an email: send it to a real inbox you control, open it, and click every link.
  • For a data change: query the actual rows and show the before and after.
  • For a fix: show the original problem no longer happens, and check whether anything else had the same problem.

Then ask for the evidence in the report itself. Not “the email was sent,” but “here’s the delivered email, and here’s what each link returned when clicked.” A report that neither includes the evidence nor points to it is an assertion, not a record.

Where adversarial review fits

A separate reviewer is especially good at catching claims dressed as evidence, because it didn’t watch the success messages roll by. When it reads a summary that says “verified the email,” its job is to ask: verified how? Where’s the proof? If the answer is “the send command returned no errors,” that’s a finding.

You can build this directly into how a reviewer works. Our free review skill includes a rule that says “it says it works” is not evidence. Tests passing, a success message, or a comment that says “done” prove nothing on their own.

It sounds pedantic. It’s the difference between finding the dead link before your users do and finding it after.

The small version you can start today

Next time your AI says something is done, ask one follow-up question:

“Show me the thing itself, not the message that says it worked.”

If it can, you’ve lost a few seconds. If it can’t, you’ve found something worth checking before a customer trips over it.

For the full approach, see our plain-English guide to adversarial review.