What replaces velocity when agents write code
Story points measured effort, and an agent can emit a hundred in an hour. What survives is verified throughput, first-pass rate and where the time went.
Velocity measured effort, and effort stopped being the scarce thing. Replace it with three numbers that still mean something when the worker is a machine: verified throughput (completions that passed a gate, never raw completes), first-pass verification rate (what share passed on attempt one), and cycle time split by phase (where the time actually went). None of them can be inflated by producing more output faster, which is precisely why they survive.
Why did velocity break rather than just drift?
Because it was a proxy for something that changed by an order of magnitude while the thing it was proxying for did not.
Story points were always an estimate of effort, and effort was a reasonable stand-in for value because producing work was the expensive part. That relationship is what broke. An agent can generate the code for a story in seconds, so the effort estimate no longer describes anything scarce. A team whose agents produce a hundred points in an hour has not delivered a hundred points of value; it has demonstrated that the unit stopped measuring.
The industry data shows the split arriving in exactly the shape you would predict. Teams with heavy AI adoption have been observed merging far more pull requests while review time rose sharply, with net productivity gains landing in the low single digits to around ten percent. Throughput and delivery came apart. DORA's 2025 report found the same tension from the other direction: AI adoption improving throughput while relating negatively to delivery stability.
Velocity cannot represent that. It goes up in both the good case and the bad one.
Should you keep estimating at all?
Yes, but for planning, and stop reporting it as an outcome.
An estimate is still useful for the thing estimates were originally for: deciding what fits, spotting a task that is much bigger than it looks, and prompting the conversation where somebody says "that is not a two, that is a rewrite". That conversation has value even when the number is wrong.
What has to stop is the second job estimates acquired, where the sum of them became a performance measure. That was always a bit dishonest and is now unworkable, because the same estimate can be satisfied by four seconds of machine time or two days of human time, and the number cannot tell you which.
One practical note if you do keep estimating: use a scale with real gaps, and have the field a person reads be the same field the chart counts. T-shirt sizes plus a lookup table elsewhere means two representations of one fact, and they drift.
What is verified throughput?
Completions that passed verification, per period. The word doing the work is verified.
Raw completion counts are the vanity metric of agent-era engineering, and they are worse than useless because they look exactly like progress. A worker that can mark forty tasks done in an hour turns "completions per week" into a measure of how fast your workers can assert. The chart rises whether or not anything was delivered, and somebody will present it.
Verified throughput is the same chart with the numerator fixed. It moves more slowly. It is occasionally embarrassing. It is the only version that survives a stakeholder who checks, and the gap between the two lines is itself the most informative thing on the page: it is the volume of work that was claimed and did not hold up.
Which numbers should actually be on the wall?
Four, and they answer different questions:
| Metric | Question it answers | Why it survives agents |
|---|---|---|
| Verified throughput | How much really landed? | Cannot be inflated by asserting |
| First-pass verification rate | How good were the instructions? | Measures the spec, not the worker |
| Cycle time by phase | Where did the time go? | Shows time relocating, not vanishing |
| Reopen rate | Did it stay done? | Catches a gate grading itself kindly |
The pairs matter more than the singles. Throughput rising while first-pass rate falls means you are shipping specification debt. First-pass rate at 100% usually means the gate is not checking anything. Cycle time flat while build time collapses means the constraint moved into spec or review, which is the single most common finding when a team adds agents and wonders why nothing got faster.
What about the review time that ate the gains?
Measure it separately, because it is the thing you can actually fix.
Human wait share, the percentage of cycle time spent waiting on a person rather than on work, is where the AI productivity gains have been going. It is not review itself that is expensive so much as the queue in front of it: the gap between work being ready and somebody picking it up.
That is worth isolating precisely because it responds to different levers than everything else on this list. Specification quality is an engineering and writing problem. Wait share is a scheduling problem, and scheduling problems are cheap to fix once somebody can see the number.
How do you move from one to the other?
Run both for a quarter, and let the divergence make the argument.
Do not announce that velocity is cancelled. Add verified throughput beside it and wait. Within a few sprints the two lines separate, and the gap between them is the conversation: this is what we counted, this is what held up. That is a much easier case to make than an abstract one about proxy metrics, and it arrives with your own data rather than someone else's report.
Take the baseline before the agents scale up if you still can. Four to eight weeks of it beats any retrospective model, and once agent volume is high the pre-agent comparison is gone for good.
In Cognibl every headline number is computed from the verifier's records rather than from activity, which is why the figures there are smaller and duller than a usage dashboard's, and why they can be shown to somebody who checks.
Keep reading
- Where to put a human checkpoint on an agentFull autonomy is the wrong goal. Let an agent run through reversible steps and stop it before anything that cannot be undone: sending, deleting, charging.
- What is AGENTS.md, and what belongs in it?AGENTS.md is a README for agents: the build commands, conventions and boundaries an agent needs. Over 60,000 repositories ship one. What to put in yours.
- How much agent work should a human review?Reviewing everything stops being possible long before anyone decides to stop. Sampling deliberately, at a rate set by evidence, beats sampling by accident.