Skip to content
All writing

First-pass verification rate, explained

Aarav Pundir · FounderMetrics4 min

First-pass verification rate is the share of agent tasks that pass on attempt one. It measures the quality of your specification, not the quality of your agent.

First-pass verification rate is the percentage of agent tasks that pass verification on their first attempt, with no rework and no human correction. It is the single most useful number to put on an agent programme, and the reason is counterintuitive: it tells you far more about how well you wrote the task than about how well the agent did it.

What exactly is being counted?

One task, one attempt, one verdict. An agent picks up a task, does the work, attaches its evidence, and the gate either accepts it or refuses it. If it passed the first time it entered verification, it counts. If it came back a second time, it does not, however good the second attempt was.

Two boundaries matter and both are easy to get wrong. A task that was never verified at all does not belong in the denominator, or you are measuring your coverage rather than your quality. And a task the agent abandoned before attaching anything is a failure, not an absence: excluding it flatters the number in exactly the case you most want to see.

Why does it measure the specification rather than the agent?

Because an agent that fails verification usually did the thing it was asked to do. It just was not asked for the thing you wanted.

A person given a vague task will ask a question, or guess and mention the guess in stand-up. A model does neither. It resolves the ambiguity silently, in whatever direction its context makes most likely, and returns something confident. When the gate then refuses it, the refusal is nearly always tracing back to a definition of done that did not say what "done" meant.

This is why the metric behaves the way it does in practice. Swap the model and the rate barely moves. Rewrite the definitions of done and it moves a lot. The industry has converged on the same conclusion from the other direction: Redis's 2026 context engineering survey found that 73% of practitioners say agents fail more often from broken context than from broken models, and 83% say fresh context matters more than more parameters.

How does it relate to change failure rate?

It is the same idea moved earlier in time.

DORA's change failure rate asks what share of deployments cause a problem in production. It is a post-release measure, and it is excellent, but by the time it moves the bad change has already shipped. First-pass verification rate asks what share of work is right before anything lands. Same question, asked at the gate instead of at the incident.

That earlier position matters more than it used to. The 2025 DORA report found AI adoption improving throughput while having a negative relationship with delivery stability: more change arriving, more of it failing. A metric that only fires after release is measuring a queue that is getting longer.

What counts as a good rate?

We are not going to invent a benchmark. Nobody has a defensible cross-industry figure for this yet, and a number made up to sound authoritative would be worse than no number, because you would plan against it.

What is safe to say is how to read your own. The absolute value depends almost entirely on how strict your gates are, so it is not comparable between two companies and barely comparable between two projects. The direction is the signal. A rate falling while throughput rises means you are shipping specification debt. A rate at 100% usually means the gate is not checking anything, which is the failure mode nobody reports because it looks like success.

Read it beside cycle time, always. A high first-pass rate bought with three weeks of specification per task is not a win, and the pair catches that where either number alone does not.

How do you actually raise it?

By changing the input, since that is where the variance is.

  • Write the done-test before the work, not after. A definition of done authored once the output exists is a description of the output.
  • Make it checkable by something that cannot be persuaded. "Handles errors gracefully" is a sentence. "Returns 429 with a Retry-After header when the rate limit is hit" is a test.
  • Fix the definition, not the task, when something fails. A task that fails twice on the same ambiguity is telling you about every future task that shares its template.
  • Watch it per template, not just per project. The average hides the one kind of work that is failing constantly.

What does it not tell you?

Whether the work was worth doing. A task can pass its gate perfectly and deliver nothing anybody needed, and this metric will applaud. It measures conformance to a specification, so it inherits every blind spot the specification has.

It also says nothing about escaped defects. Verification catches what it was built to catch; the things it was not built to catch reach production with a passing report attached. That is why reopen rate sits beside it as a headline metric rather than underneath one, and why a first-pass rate presented on its own should make you suspicious of whoever is presenting it.

In Cognibl the rate is computed per project from the gate's own verdicts, not from anybody's self-report, because a pass rate calculated from claims is a survey rather than a measurement.