The Agent’s Tests Agreed With the Agent

“The verification is probably the single most important thing that people do not get right, largely.” That is Boris Cherny, who built Claude Code, talking to YC this summer. His advice: give the model a way to verify its own work, and it will just go.

Our global instructions to every agent say the opposite, in capitals: “NEVER include tests unless explicitly requested.” So we tested it. We gave the same agent the same real tasks twice. Once with our rule. Once with “before implementing, write an executable check for this task’s exit criteria, run it, and keep working until it passes.”

The rule stays. But not for the reason we wrote it, and that is the interesting part.

The setup

Two real changes from one of our repositories, replayed from the commit before anyone reviewed them:

  • Harden a benchmark report script that our reviewer had later caught printing wrong numbers.
  • Add a secret redactor to a hook that saves session memory.

The task text described the goal and the stakes, not the known defects. Then three things scored each run, none of which the agent could see:

  • Hidden fixtures. Five broken inputs for the report script (an errored cell, an unscored one, an empty answer that still carries a score, a duplicate, a missing question), and fifteen credential shapes plus eight lookalikes that must pass through untouched for the redactor.
  • A fresh adversarial review of the diff by a different model from a different provider.
  • Cost.

Five runs per variant, per task, on one model.

The result

no testswrite a check first
report script: broken inputs refused, of 50.40.2
report script: high-severity review findings4.44.4
redactor: credential shapes caught, of 1511.610.8
redactor: high-severity review findings5.84.0
cost per run, report / redactor$0.42 / $0.15$0.62 / $0.20

No measurable difference in the defects that got through. The redactor’s review count leans toward the check-first runs, but the confidence interval includes zero. The cost is 33 to 48 percent higher.

For scale: after our reviewer had been through it, the real report script refuses four of the five broken inputs. Neither variant got close.

Why

The check-first runs did write tests. The tests passed. They covered exactly the cases the agent had already thought of, which is to say they checked the agent’s reading of the task against the agent’s reading of the task. The defects that mattered lived outside that reading: credential shapes nobody listed, errored rows treated as a warning instead of a failure.

Quis custodiet ipsos custodes: who guards the guards. Juvenal meant something else, but it fits. A test written by the code’s author is a second opinion from the same person.

Now reread Boris’s examples. The Bun runtime was rewritten from Zig to Rust against test suites the agent did not write: “There’s a big test suite in Bun, there’s a big test suite in Node.js. So it’s easy to know if you did the right thing.” The agent porting a desktop app to Swift screenshots the original app and compares the two “pixel by pixel”. In both, the check is something the agent did not write and cannot talk into agreeing. That is not self-verification. It is an independent oracle that happens to be automated.

Ours is a reviewer from another provider, run adversarially against every diff. In this experiment it was the only thing that found the defects: four to six high-severity ones per diff on average, in both variants.

The same thing, in production

After the experiment, we shipped the fix for the redactor gaps it had turned up (voitta-rag#65, released as v0.1.5). Eight review rounds. Findings per round: 7, 7, 5, 3, 5, 1, 1, 2.

Before each of the nine pushes, the agent’s own regression suite was green: the fixture, every earlier round’s reproductions, and later a fuzz property. It was not useless. Twice it caught a real problem before the push. What it never did was anticipate the next round. Each time, the reviewer found shapes the suite did not contain.

The count did not come down because the tests got better. It came down when the design changed so that it could not regress against the shipped version by construction. That is part 4 of this series.

What we kept

The no-tests rule stays, on evidence this time: asking the agent to write its own check cost more and caught nothing extra. Independent review stays, and gets leaned on harder. Boris is right that verification is the thing. The word doing the work is the one he did not have to say: independent.

Limits, plainly: five runs per cell, two tasks, one model, one repository. This is a result about our setup, not a law.

Part 1 of a series on what our process experiments found. Next: the eval that read the answer key.

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.