The Tests Were Green. They Were Lying.

Six ways an AI writes a test that proves nothing — and what I learned catching them in my own system.

Alex Hinojosa

I build with AI agents. They write the code, they write the tests for the code, and the tests pass. For a while that felt like the whole promise delivered: green tests, say what you want, get working software with a green checkmark on top.

Then I started pulling on the checkmarks.

Here’s the one that broke me. An agent had written a guard to keep a contaminated year of data out of a model’s training set — an important rule; getting it wrong would quietly poison everything downstream. There was a test protecting it, and the test was green. So I read the test. It searched the source file for the string "2020" and passed if it found it.

The string "2020" was sitting in the file’s comments.

You could delete the actual filter — the real logic that did the work — and the test stayed green. It had been guarding nothing for months. Not because anyone was lazy. The test looked exactly like a test. It ran, it asserted, it passed. It just never once checked the thing it was named after.

That sent me looking, and what I found is the reason I’m writing this. I have now caught six distinct ways a test can be green and be lying — all of them in my own system, all written by capable agents, several of them in a single week. Everyone is building with agents now. Everyone is trusting the green. Almost nobody has looked underneath it. So I’ll go first.

The lesson under the lesson

I already knew the first rule of this: a rule with no verifier is a wish. If you say “we never do X” and nothing actually checks for X, you don’t have a rule, you have a hope. I’d learned that one the normal way — by getting burned.

What the green tests taught me is the level below it, and it’s worse:

A check that cannot fail is more dangerous than no check at all.

No check leaves you honestly uncertain — at least you know you don’t know. A check that cannot fail makes you confidently wrong, silently, indefinitely. It isn’t a missing guardrail. It’s a guardrail painted onto a cliff. And an AI will build you these all day, because a test that always passes looks identical — in every report you will ever read — to a test that works.

The six ways a green test lies

Every one of these is real. I’ve stripped the specifics; the shape is exact.

1. It compares a value to itself. A check meant to confirm two records referred to the same thing — by comparing each field of the record to itself. a == a. It can only pass. The detail I can’t get past: an agent wrote this while fixing a ticket about checks that can’t fail. It had the concept in its context and produced a fresh instance of it in the same breath.

2. It reads the source code instead of the behavior. The "2020" story above. The test greps the file for a string instead of running the thing and watching what it does. Any check that inspects code as text rather than exercising it as behavior is one stray comment away from meaningless.

3. It passes because the language let it. A function’s inputs changed — it now expected an ID and a category. The tests still called it the old way, with a date and a dictionary. The language shrugged, accepted the wrong types, quietly turned the dictionary into a string, and every test passed without ever touching a real ID. No single assertion was wrong. The test just never reached the code that mattered.

4. It tests a copy of the logic, not the logic. Five tests for one production guard. Not one of them imported the guard. Each had quietly redefined its own little copy of the function inside the test file and tested that. You could delete the entire production guard — the real one, the one that ships — and the suite stayed green, cheerfully validating a duplicate that no longer existed anywhere real.

5. It gives up quietly when it can’t check. A guard that needed to load a reference index to do its job. If the index failed to load, the guard didn’t raise an error — it just returned, and the unverified value sailed through and got written. “I couldn’t check this” and “I checked this and it’s fine” produced the identical outcome: nothing happened. Cannot-verify is not verified.

6. It stops the bleeding and leaves the blood. A fix that repaired the thing producing bad data — and left all the bad data already written sitting exactly where it was. Or its mirror: someone scrubs the bad data but never lands the guard that stopped it being produced. That’s not a fix. It’s a countdown.

Why AI makes this so much worse

Humans write bad tests too, of course. But agents fail this way structurally, and it’s worth being precise about why, because it changes how you defend against it.

An agent starts every task as a fresh mind. It didn’t sit through the outage. It wasn’t on the thread where you all swore never to do the thing again. It knows what’s in its prompt and the files it loads, and it knows nothing else — not one thing. So when an agent wrote that assertion comparing a value to itself while fixing a bug about exactly that, it wasn’t being careless. It had never been told. And it couldn’t have been, because the discipline lived where most hard-won engineering wisdom lives: in people’s heads, and in old threads nobody reloads.

That’s the real lesson, and it’s the one that transfers to anyone building with these tools:

If a discipline isn’t written in a file the agent loads, it does not exist. And every agent is a fresh mind, eventually.

The most important rule in your project is worthless as tribal knowledge. Humans absorb tribal knowledge by osmosis, slowly, over months of standups and mistakes. An agent absorbs exactly what you hand it and starts from zero every single time. You have to write the wisdom down — in the file it reads before it writes a line — or it will confidently reinvent every mistake you have already made.

How you actually catch them

Three things moved the needle, and none of them is “write more tests.”

Give something one job: try to fool the test. I built a reviewer whose only question is “can I make the underlying thing wrong and still get green?” Not “does this look correct” — correct-looking is how all six shipped. It works by mutation, not argument: it breaks the real thing on purpose and watches whether the check turns red. If the check stays green while the world is broken, the check is a lie, full stop. No debating whether it would have caught it.

Prove the verifier fires. This is the whole game in one habit. Break the thing on purpose and watch the test go red with your own eyes before you ever trust it green. A test you have never seen fail is not a passing test — it is an untested one. An empty result and a broken instrument look identical until you feed it something it is supposed to reject.

Instrument the thing; don’t reason about the thing. My most expensive bugs came from confidently reasoning about what a system ought to do. The fix, every time, started with logging what it actually did. One guard I had to rewrite three times; the only version that worked began with me dumping the raw input to a file and reading it, instead of trusting my mental model of it. Your mental model is where the lies hide. The log doesn’t lie.

Go first

Here’s the discipline that ties it together, and it’s the sequel to the rule I opened with:

Every rule ships with its verifier. And every verifier ships with its own counterexample — the broken case you have personally watched it catch.

Otherwise you’ve just moved the wish down one level and stopped looking at it.

I’m laying out all six of my own failures in public because the alternative — everyone quietly shipping green suites that prove nothing, and calling it working software — is already happening, at scale, right now. The green checkmark has never been cheaper to produce or easier to fake, and we have handed the pen to something that fakes it fluently and without a trace of malice.

So look under yours. Break one thing on purpose today and see if anything turns red.

And the line I would put on the CI dashboard if I could:

A check nobody has tried to fool has not been tested. It has only been admired.


If you want the full treatment — the disciplines these green-test failures became — start with the field guide (especially Part 1: checks that can fail) and the companion playbook (§6 hooks).

0 replies

Leave a Reply

Want to join the discussion?
Feel free to contribute!

Leave a Reply

Your email address will not be published. Required fields are marked *