Skip to content
All writing

How much agent work should a human review?

Aarav Pundir · FounderReview5 min

Reviewing everything stops being possible long before anyone decides to stop. Sampling deliberately, at a rate set by evidence, beats sampling by accident.

Not all of it, and the honest question is not what percentage but what decides the percentage. Review everything and the queue becomes the bottleneck within weeks. Review nothing and you have no idea what your gates are missing. What works is a rate tied to evidence: high while a worker is unproven, falling as it earns it, and rising again the moment the numbers say something changed.

Why does "review everything" stop working?

Because it stops being true before anybody decides to stop.

The transition is quiet. A team reviews every completed task, agent volume climbs, the reviews get faster, then shallower, then become a scan of the summary. Nobody announces a policy change. The stated process still says everything is reviewed, and it is now false, which is worse than an honest sampling policy because the gap is invisible to everyone including the people doing it.

The measurable version of this shows up as human wait share: the fraction of cycle time spent waiting on a person. Once agents are producing work faster than people can absorb it, that share grows without limit, and the queue converts directly into cycle time. You are then paying the full cost of review while getting a progressively smaller amount of actual reviewing.

So the choice is not between sampling and not sampling. It is between sampling on purpose and sampling by accident.

What should set the rate?

Evidence about that specific worker, applied by a rule rather than by mood.

The inputs that earn their place:

  • Track record on this kind of task. A first-pass rate measured over enough work is the strongest single signal, and it is specific: an agent can be excellent at one task template and poor at another.
  • How new the worker is. Everything unproven starts at a high rate. A new agent, a new model version, a new task type. Probation is not an insult, it is how you get data.
  • What the task can reach. Irreversible or externally visible work carries a higher rate regardless of track record, because the cost of the miss is different, not because the worker is worse.
  • Whether anything changed. A model upgrade, a prompt change, a new tool. Any of these resets what your history is evidence about.

What should not set it: how busy the team is, how much somebody trusts a particular vendor, or whether the release is on Friday. If the rate moves for those reasons it is not a control, and the first time it matters will be the first time it was lowered.

Should the rate be the same for everything?

No, and making it uniform is the most expensive mistake available.

A flat 10% across all work spends the same attention on a documentation fix and a change to billing. The whole value of sampling is that it lets you concentrate scarce review where the loss is large, and a uniform rate deliberately discards that.

Stratify on consequence first and confidence second. Work that touches money, customer-visible output, permissions or data deletion sits in a high band permanently. Routine internal work with a long clean record sits low. The middle, which is most of it, floats on measured performance.

The useful sanity check: if the rate never differs between two kinds of task in your system, you have a policy rather than a sampling strategy.

What should a reviewer actually look at?

The evidence, and specifically the places where evidence and claim diverge.

Handing somebody a completed task and asking whether it is good is the least effective possible use of the review. They will read the summary, find it plausible, and approve. Plausibility is what the worker is best at producing, so plausibility is the wrong thing to test.

What makes a review productive is arriving with the mechanical checks already done and the disagreements surfaced: the deterministic checks that ran and what they said, the definition of done beside the attached evidence, and any point where the two do not line up. Then the human is deciding the genuinely judgemental question rather than re-performing work a machine already did.

This is also what makes a sampled review worth more than an unsampled one. If every review is a full manual re-verification, you can afford very few. If the review is adjudication on top of automated checking, you can afford enough of them to learn something.

What do you do with what the sample finds?

Change the gate, not just the task.

A sampled review that catches a problem has told you two things: this task was wrong, and your automated checks did not catch it. The second is much more valuable and is usually ignored under time pressure. Every escape found by sampling is a specification for a check that should have existed.

The rate should respond too. A failure found in a sample is evidence that this worker's rate was set too low, and the adjustment should be automatic rather than a debate.

The measure that tells you whether the whole system is working is auto verification share: the proportion of work that passed with no human touch at all. You want it rising over time, because that means the gates are absorbing more of the load, and you want it rising alongside a stable reopen rate. Rising auto-verification with a rising reopen rate means the gates got more permissive, not better.

Where should you start?

High, and come down with evidence.

Starting low and raising it on the first incident is the wrong direction, because the incident is the thing you were trying to avoid, and because the rate you set on a bad day tends to become permanent. Start by reviewing a large share of a new agent's work, watch the first-pass rate stabilise, and let the rate fall as the record accumulates.

Write down what would push it back up before you need it. A rate that only ever falls is not a control either.

In Cognibl the human sample sits behind the deterministic checks and, on the AI plan, behind a flow that reads the proof against the definition of done. The sampled share is set by the agent's score rather than by anyone's mood, and a manual override is possible but requires a logged reason and shows in the audit trail.