Adversarial Review

A check that can't fail is not a check

Many projects built with AI end up with automated checks. Every time the AI proposes a change, a robot runs the tests, and a green tick or a red cross appears next to it. Green means go.

It is reassuring. It is also easy to quietly break.

The temporary setting that stayed

On one software project, a new set of automated tests was being wired up. While it was still settling in, the step that ran those tests was marked as non-blocking: it would run and report, but a failure would not stop anything. The setting even had a note next to it, saying that once the tests had been green for a while, a follow-up change would make them blocking again.

The follow-up never happened.

About two weeks later, someone looked into why the test step was showing red. The investigation found two problems that had landed in the main version of the code and stayed there, because nothing could stop them.

The first was a test whose job was to make sure sensitive figures were never shown on screen without protection. A perfectly good fix elsewhere had changed how one of those figures was displayed, in a form the check did not recognize. The test went red that day. Nobody noticed, because red no longer meant anything.

The second was stranger. A new test had been added that had never passed, not once, from the moment it was written. It was written in a format the test runner could not even load. Every run since, it had failed. Every run since, the overall result had still been allowed through.

Nothing was hidden. But a note is not a reminder, and an AI session that starts without the earlier context will not remember what was promised for later.

The same shape, without the setting

A related problem turned up on the same project in a different form. An important protection existed in the code: a rule that stopped one user from editing another user’s profile. The rule worked. But no test covered it, so if a later change had deleted it by accident, no check would have flagged it.

When a test was finally written for it, the work was checked one more way: the protection was temporarily removed and the tests were run again. The tests that should have failed did fail, and the unrelated ones stayed green. Only then was the test trusted.

That extra step is the habit this post is about.

Why this keeps happening with AI-built software

Automated checks have two jobs. Saying yes when things are fine is the easy one. Saying no when things are broken is the one that matters, and it is the one you almost never get to see.

A check that is switched to “report only” can show red and still let everything through. A test that checks the wrong thing, or a test that is missing entirely, can leave the overall result green while the thing you care about is unprotected. You cannot tell a working alarm from a disconnected one by how quiet it is.

AI coding tools make this more likely, in practice, for a few reasons. They are good at getting things to pass, and “make it non-blocking for now” is a quick way to get past a red cross. They often write the tests and the code together, so the test may have only ever seen code that already worked. And unless earlier intentions are written where the next session will look, a “tighten this later” note depends on someone remembering it. (The wider version of this problem is covered in The author can’t grade the paper.)

The habit: make every check fail once, on purpose

Before you trust a safety check, watch it say no.

You do not need to read code to do this. Ask your AI assistant, in plain words:

  1. “List every automated check that runs on a change, and tell me which ones can actually block it.” Ask specifically whether any are set to continue on failure, marked optional, or not required. Anything that cannot block is a report, not a gate. Decide on purpose which is which.
  2. “For the most important protection in this feature, temporarily break it on a separate test copy of the code, run the checks, and show me which ones fail. Then put it back.” Do this away from your live product, never where real users are. You want to see a red result that names the thing you care about. If everything stays green with the protection removed, the checks that ran did not catch its removal. Maybe there is no test, maybe the test is skipped, maybe it checks the wrong thing. Either way, you cannot count on it yet.
  3. “Show me any test that has been failing for a long time, and whether it has ever passed.” This depends on the project keeping a history of past runs, so the answer may be incomplete. A long-red test is either broken itself or pointing at a real problem nobody has fixed. Find out which. In the story above it was broken, and because failures were allowed through, its redness taught everyone to ignore red.

One deliberate break, on the one or two things that would hurt most if they failed, goes a long way toward showing whether your alarm is wired up for the failures you fear most.

How to check it worked

The proof is a screenshot or a pasted result you have actually looked at, in two parts:

  • a run where the protection was removed and the check went red, naming the right thing, and
  • a run after it was restored where the check went green again.

If your assistant tells you “the test would catch that,” ask to see the red run. That sentence is a claim, not evidence.

And if you ever switch a check to “report only” while setting something up, do not rely on a note. Put the switch-back on your own task list, with a date, the same day you switch it off. A temporary exception with no owner tends to become the permanent rule.

If you want a second pair of eyes on a change before it goes in, our free review skill asks the reviewer whether any tests were deleted, skipped, or loosened to make things pass.