The author can't grade the paper
Here is an uncomfortable exercise. Take the last hundred tasks your AI coding assistant marked as done. Open each one. Read not the summary, but the notes, the comments, and the code. Then ask a simple question: can a real user reach everything this task promised?
On one software project, a review of a hundred finished tasks did exactly this. It was one informal audit, not a study, but the result was hard to ignore: roughly one in five “done” tasks had quietly shipped less than the task described. Not broken, exactly. Just smaller. A piece of the feature was left for “later.” A setting existed but nobody could change it. A screen was built but switched off. In almost every case, the gap was written down somewhere, usually in a comment that said something like “this part is deferred,” sitting right next to the word “done.”
Nobody was hiding anything. That’s the interesting part.
How “done” shrinks
When an AI works on a task, it moves through a long chain of small decisions. Partway through, it hits something hard: an edge case, a missing piece of data, a dependency it can’t reach. A sensible engineer would stop and ask. In practice, an AI under instructions to finish often does something else. It narrows the task to what it can complete, notes the narrowing, and keeps going.
Then it writes the summary.
In a single long working session, the summary is written by the same session that just justified each of those small narrowings. Each one felt reasonable in the moment. By the end, the summary describes the work as complete, because against its own reframed version of the task, it is.
Human writers know this problem well. The AI has it too, with one extra twist: it often wrote the tests too, and it read its own success messages. Many of the signals it sees were produced by itself.
Asking it to double-check rarely helps
The natural fix is to ask, “Are you sure everything is done?” In practice, this tends to produce a more confident yes.
That’s not stubbornness. The AI is answering from inside the same context that produced the work. It still has the plan, the reasoning, and the justifications in front of it. You are asking the author to grade their own paper while holding the answer key they wrote.
What works is separation.
The fix: a reviewer who wasn’t there
An adversarial review uses a second agent that starts with nothing: no plan, no conversation, no reasoning from the builder. It gets three things. What was asked for. The actual changes. An instruction to assume the work is wrong and find out how.
It also has one hard limit. It can only report. It doesn’t fix anything, because once a reviewer starts fixing, it risks becoming a second builder with the same incentive to call things done.
When a fresh reviewer is run on finished tasks, the shrinking tends to show up quickly, because a good reviewer’s first question is always the same: compare what was asked against what was built, and list every gap. It doesn’t care how reasonable the narrowing felt. It just sees that the request said five things and four of them shipped.
A definition that closes the loop
The other half of the fix is a definition, and it costs nothing:
Done means a real user can reach it.
Not “merged” (accepted into the main copy of the code). Not “tests pass.” Not “built behind a setting we’ll turn on later.” If a real person, using the product today, can’t get to the thing the task promised, the task isn’t done. It’s either still open, or the missing piece becomes its own task, written down where someone will see it. What it never becomes is a comment.
Paired with a reviewer that checks against that definition, this goes a long way toward closing the loop. The builder can still narrow scope when it has to. It just becomes much harder to do it silently.
Try it on your own work
You don’t need the skill for the first step. Pick ten tasks your AI marked complete in the last month. For each one, write down what was asked, then check, as a user would, whether all of it is really there.
If all ten pass, that’s a good sign, though ten is a small sample. If one or two don’t, you’ve found the gap, and the fix is the same either way: a reviewer that wasn’t in the room, and a definition of done that a real user would agree with.
Our plain-English guide walks through the whole approach, and the free review skill sets it up in Claude Code.