Skip to content
All writing

Cycle time for agent work: the four phases

Aarav Pundir · FounderMetrics4 min

One cycle-time number hides where the time actually went. Split it into spec, build, verify and settle, and the bottleneck in agent work becomes obvious.

Cycle time for agent work should be reported in four parts, not one: spec (the task was created until the first run started), build (the run itself), verify (evidence attached until the gate returned a verdict) and settle (passed until closed). A single nine-day number tells you nothing you can act on. The split tells you which of the nine days to go and fix.

Why does one number stop working once agents are involved?

Because agents move the bottleneck without moving the total.

With a human team, execution was the expensive phase, and everyone knew it. Most tracker metrics were designed around that assumption: time in progress, work in progress limits, burndown against days of effort. They all quietly encode the belief that building is the slow part.

Agents break the assumption. Execution collapses toward minutes, and the time that used to hide behind it becomes the whole story. Teams then report a cycle time that has barely improved and conclude the agents are not working, when what has actually happened is that the constraint moved into a phase nobody was measuring.

DORA found the same shape at the release end. The 2025 report recorded AI improving throughput while delivery stability declined: more work arriving faster, more of it needing attention afterwards. Time is not saved so much as relocated, and an undivided metric cannot see relocation at all.

What are the four phases?

Spec runs from creation to the first run. It holds everything that has to exist before an agent can start: the description, the definition of done, the scope, the revisions to all three. It is the phase most teams do not know they have.

Build is the execution itself, whoever does it. For an agent task this is often the shortest phase on the chart, which is exactly the finding.

Verify runs from evidence attached to verdict returned. It covers the automated checks, any model-graded checks, and the wait for a human when the task was sampled. It is the phase most likely to be pure queue.

Settle runs from a passing verdict to a closed task: the merge, the release, the sign-off, whatever your process still requires after the work is correct.

Which phase is usually the problem?

Spec and verify, and they fail in opposite ways.

Long spec time is not automatically bad. Writing a precise done-test is real work, and it is the work that makes the run succeed first time. Long spec with a high first-pass rate is a team investing correctly. Long spec with a low first-pass rate is a team rewriting the same task repeatedly, which is a different problem wearing the same timestamp.

Verify time is more often waste, because most of it is not verification. It is the gap between the evidence being ready and a person picking it up. That gap belongs in its own measure, human wait share, and it is usually the largest single lever available. It also responds to scheduling rather than to engineering, which makes it the cheapest thing on this list to fix.

How do you instrument the split honestly?

From status transitions, and only from transitions the system itself recorded.

Each phase boundary has to be an event the tracker wrote when it happened, not a field somebody set afterwards. The moment a phase boundary becomes editable, the decomposition becomes a story about how the work is supposed to go. Three rules keep it honest:

  • Take the median, not the mean. One task that sat over a holiday will dominate an average and tell you nothing about the normal case.
  • Count reopens as new cycles, not extensions. Otherwise a reopened task silently inflates settle time and the reopen itself disappears.
  • Report the unestimated and the unverified as counts beside the chart, never folded in. A number that quietly excludes the awkward cases is the number people stop trusting first.

What do you do with the answer?

Each phase has its own fix, and they are not interchangeable:

Phase dominatesThe actual problemWhat moves it
Spec, with low first-passAmbiguous definitions of doneTemplates, checkable criteria
Spec, with high first-passUsually nothingLeave it alone
BuildGenuine execution difficultyScope, tooling, model choice
VerifyWaiting on a humanReview scheduling, sampling rate
SettleProcess after the work is rightRelease cadence, approvals

The reason to split before acting is that three of those five rows look identical on an undivided chart. A team that adds agents to fix a spec problem has bought a faster way to produce work nobody specified.

Is this just lead time with extra steps?

No, and the difference is the verify phase.

Lead time and cycle time as normally defined stop at delivered. They have no concept of a verdict, because when people did the work the claim of completion and the completion were the same event. Once an agent is the worker those come apart: the claim arrives instantly and always, and the verdict arrives later and sometimes says no. A metric with no verify phase cannot represent the gap between a task saying it is done and a task being done, which on an agent programme is the gap that matters most.

In Cognibl the decomposition is the headline chart on a project, drawn as one stacked bar per phase, because the comparison people need is between the four parts rather than against last quarter.