Each change passed. The combination never ran.
A common routine for software built with AI: the AI proposes a change, an automated system runs the tests, a green tick appears, and the change is merged into the main version of the code, the one your live product is built from. Many projects let that merge happen by itself once the tick is green. It feels safe, because nothing goes in without passing.
There is a gap in that reasoning, and on some setups it is wide open.
What the green tick tested
The tests on a proposed change run on a particular version of the code. Depending on how a project is set up, that can be the change on top of main as it stood when the tests ran. If two changes are open at the same time, each can be tested without the other, both can pass, and once both are merged, the code on main is a combination no test has seen.
Some setups close this gap before merging. GitHub can require each change to be brought up to date with main before it merges, and it offers a merge queue, covered below. The project in this story used neither, so its safety net was the usual one: the tests run again on main after every merge, and a broken combination turns up soon after.
That safety net is the part that went missing.
The days the tests on main stopped
On one software project, a small automation was added so that changes merged themselves once their tests passed. A couple of days later, a scheduled AI check that watches the project’s test runs noticed something: since the automation was switched on, no merge had started a test run on main. A series of changes had gone in, each green on its own, and none of the merges had been followed by the tests that normally ran on main.
The cause was a single setting. The automation merged changes using the default credential that GitHub provides inside its automation service, GitHub Actions. GitHub’s documentation says that events caused by that default credential, with a couple of narrow exceptions, do not start new workflow runs. It is a deliberate guard against automations triggering each other in an endless loop. On that project, the side effect was that those merges started no tests on main. No error, no warning, no red cross.
Two of the merged changes had modified the same part of the product, each tested without the other. When main was finally built and tested by hand, the tests that are meant to block a release passed, so as far as anyone could tell nothing had broken. That later check could not show that each merge had been safe at the time it happened. It only showed that the code was fine on the day someone looked.
The fix needed a person: a separate credential for the automation, so GitHub treats its merges like ordinary updates and starts the tests. Creating a credential is a decision about access and secrets, so it went to the project’s owner. Until then, the scheduled check took on a stopgap: build and test main by hand whenever new merges arrived.
Why this is easy to miss
This is a different problem from a check that runs but cannot fail (covered in A check that can’t fail is not a check). Here the tests could fail perfectly well. They were just not being run.
Every change you look at has a green tick. What is missing is a test run on something you were not looking at, and a run that never happened does not show up as a failure anywhere. In practice, an AI asked to “make changes merge automatically” can reach for the simplest setup, which uses the default credential. It does what was asked, so the session reports success.
The habit: match merges to test runs
Once a week, and right after you turn on any automation that merges code, line up two lists and compare them: the merges into main, and the test runs on main.
You do not need to read code for this. Ask your AI assistant:
“List the last ten merges into main, with the time of each. Then list the test runs on main over the same period, with the version each one tested. Show me any merge that was not followed by a test run on that version, and give me the link to each test run.”
Then open two or three of the links yourself. A test run’s page on GitHub shows which version of the code it ran on and whether it passed. You are looking for one thing: the version it tested should be the merge it claims to cover. If the AI reports a gap, ask a second question: “List every automation that merges code here. What credential does each use, and does a merge it makes start the tests on main?”
If merges are not followed by tests, there are two reasonable fixes. One is this project’s: give the merging automation its own credential so its merges count as ordinary updates. The other is GitHub’s merge queue, which tests each change together with the changes queued ahead of it before merging. It is not available on every plan, and your tests need a small setting change to run on queued changes, which your assistant can make. Either way, creating credentials or changing who can merge is a decision about access, so it should be yours, not the AI’s.
How to check it worked
After the change, make one small, harmless merge. Then check the evidence that matches the fix you chose.
- If you gave the automation its own credential: ask for the link to the test run that followed your merge, open it, and confirm it ran on the merged version and passed.
- If you switched to a merge queue: the test happens before the merge, so a fresh run on main is not the evidence. Ask for the link to the queue’s test run for your change, open it, and confirm it finished and passed.
“The tests run on every merge now” is a claim. The evidence is a test run you opened yourself, on the version that was merged.