How to verify what an AI agent actually did
Reading every diff does not scale and spot-checking misses what matters. A three-layer method: decide by machine, judge by model, sample by human.
The first week of running agents feels fine because you read everything. The fourth week does not, because you cannot. Somewhere in between, review quietly becomes sampling, and the sampling is not deliberate: it is whatever you had time for. That is the moment the honest answer to "is this work good" becomes "some of it, probably".
The fix is not more attention. It is deciding, in advance, which questions get answered by a machine, which by a model, and which by a person, and never spending a person on a question the first two could have settled.
Layer one: what a machine can decide
Start here, because it is free and it never gets tired. A surprising share of "did the agent do this properly" reduces to questions with unambiguous answers:
- Does the test suite pass, and did coverage move the way the task said it would?
- Do the required fields of the proof of work exist and contain something?
- Does the diff stay inside the files the task named?
- Does the referenced commit or document actually exist, and is it the one the report was pinned to?
None of these needs judgment. All of them are common failure modes, particularly the last: work that references something which has since changed is one of the easier ways for a green tick to become meaningless.
Everything a deterministic checker settles is something no reviewer has to open.
Layer two: what a model can judge
The second layer handles the questions that need reading but not authority: does the evidence actually support the claim, and does it match what was asked?
This is where a model earns its place, and where it needs a constraint to be useful. Do not ask "is this good work": that invites a confident essay with no falsifiable content. Ask it to compare two specific documents: the definition of done, and the proof of work attached to the task. Then have it report where they agree, where they do not, and what it could not tell either way.
The last category is the valuable one. A model that says "the definition asked for a rejects file and I cannot find one referenced anywhere" has done something genuinely useful, and it has done it in a form a person can check in seconds.
One rule keeps this layer honest: the model reports, it does not decide. The moment its judgment becomes the gate, you have replaced one unverifiable claim with another and added a step.
Layer three: what only a person settles
Some things stay human. Whether the outcome was worth having. Whether the approach will be a problem in six months. Whether the thing that passed every check is nonetheless the wrong solution to the right problem.
The design question is not whether to sample but at what rate. Sample everything and you have not automated anything. Sample nothing and the first two layers are unaudited, which matters because a checker with a bug is a checker that passes everything.
So the rate should be mechanical rather than discretionary: derived from the agent's track record, applied by formula, and raised when the record gets worse. A new agent should be sampled heavily. One with a long clean history, less. Any manual override of that rate should be logged with a reason, because "we waived review that week" is precisely the thing you want to be able to find afterwards.
What makes any of this possible
All three layers depend on something that has to be true before verification starts: the record of what happened has to be trustworthy.
If an agent reports its own actions, you are verifying a story it told about itself. That story will be internally consistent and it may be entirely wrong, and no amount of careful reading at layer two will reveal the difference. When the tool calls pass through a gateway that writes the log, the trace is a record of what the agent did rather than what it said it did, and the two layers above have something real to work with.
Same for evidence. Proof of work that can be edited after the fact is a document about the present, not about the moment of approval. Versioned and append-only, it can answer the question that actually comes up: what did this look like when somebody said yes.
Where to start
Do not build all three at once. Take the last thing that went wrong and ask which layer would have caught it.
If a test would have caught it, you have a layer-one gap, and it is the cheapest to close. If it needed someone to read the definition of done next to the output, that is layer two. If it took judgment about whether the work was worth doing, that is layer three, and it means your sampling rate is too low or pointed at the wrong tasks.
Most teams find the first answer more often than they expect, which is good news: it is the layer that scales without asking anybody for their afternoon.
Keep reading
- Where to put a human checkpoint on an agentFull autonomy is the wrong goal. Let an agent run through reversible steps and stop it before anything that cannot be undone: sending, deleting, charging.
- What replaces velocity when agents write codeStory points measured effort, and an agent can emit a hundred in an hour. What survives is verified throughput, first-pass rate and where the time went.
- What is AGENTS.md, and what belongs in it?AGENTS.md is a README for agents: the build commands, conventions and boundaries an agent needs. Over 60,000 repositories ship one. What to put in yours.