Measuring agentic AI: what to instrument
Most teams running agents in production have no measurement framework at all. Here is the short list of what to instrument, and what to ignore.
Instrument four things and you can run an agent programme: whether the work passed on its first attempt, where its cycle time actually went, how much of that time was spent waiting on a person, and how much of it came back afterwards. Everything else is either downstream of those four or is activity data wearing a metric's clothes.
The gap this closes is real and it is wide. Arcade.dev's 2026 survey found 52% of gen-AI organisations running agents in production and only 31% with any measurement framework for them. Most teams shipped the worker before they built the instrument.
Why do the usual AI metrics not answer the question?
Because they measure the model, and your problem is the work.
Token counts, latency, tool-call frequency, hallucination rate: all worth having, all genuinely useful for debugging a run. None of them can tell you whether the task got done. An agent can produce a fast, cheap, perfectly grounded answer to the wrong question, and every one of those numbers will look excellent.
The observability tooling that has grown up around agents is mostly built for the first job. The industry's own reviews describe the split plainly: the capability that is thin is not the depth of tracing but policy enforcement at execution time and attribution good enough to stand up. You cannot attribute what you did not log, and you cannot manage what you only logged after the fact.
What are the four that matter?
First-pass verification rate. The share of agent tasks that pass on attempt one. It is the specification-quality metric, and it is the closest thing to a single health number this work has.
Cycle time, decomposed. Split into spec, build, verify and settle. Agents collapse the build phase and relocate the time rather than removing it, so an undivided number tends to show no improvement at all and gets read as failure.
Human wait share. The percentage of total cycle time spent waiting on a person: review pickup age, sample-queue age, blocked-on-owner. It is usually the largest single lever and the cheapest to pull, because it responds to scheduling rather than to engineering.
Reopen rate. Tasks reopened within thirty days of completion. This is the one that catches a verification system grading itself too generously, which no measure taken at the moment of completion ever can.
What about throughput?
Count it, but count verified completions only, never raw ones.
Raw completion counts are the vanity metric of this whole field, and agents make them meaningless overnight. A worker that can mark forty tasks done in an hour turns "completions per week" into a measure of how fast your workers can type the word done. The chart goes up and to the right regardless of whether anything was delivered, which makes it worse than no chart, because someone will present it.
Verified throughput is the same shape with the numerator fixed: completions that passed a gate. It moves slower, it is occasionally embarrassing, and it is the only version that survives contact with a stakeholder who checks.
Where should the numbers come from?
From the system that observed the work, never from the worker that did it.
This is the whole design constraint, and it is easy to get wrong because self-reported telemetry is so much simpler to build. An agent that writes its own trace is a witness testifying about itself. It is not usually lying, but it is reporting what it believed it did, which is precisely the thing in question when a task fails.
So: record tool calls at the gateway the calls pass through. Record verdicts from the checker that issued them. Record timestamps from the transitions the tracker itself wrote. If a number can be edited afterwards by the party it evaluates, it is a claim rather than a measurement, and it should be labelled as one everywhere it appears.
What should you deliberately not instrument?
- Agent signups and raw job counts. Volume, not value. They headline nothing.
- Per-run token cost as a headline. Worth watching for budget, useless as a quality signal, and it pushes teams toward cheaper runs that fail more often.
- Anything a person maintains by hand. A metric that requires somebody to keep a field accurate decays quietly, and you will not know when it stopped.
- Model-level benchmarks as a proxy for delivery. A better score on a public benchmark has never once told anybody whether their tasks are passing.
How long before the numbers mean anything?
Four to eight weeks of baseline before you draw conclusions, and take the baseline before the agents arrive if you still can. The guidance from teams doing this well is consistent on the point: instrument the workflow first, because a pre-agent baseline beats any retrospective benefits model.
The awkward part is that the first honest report usually looks worse than the optimistic one it replaces. Verified throughput is lower than raw throughput. First-pass rate on a new programme is often unflattering. That drop is not a regression; it is the difference between what was happening and what was being counted, arriving all at once.
In Cognibl these five sit on the project home page, computed from the verifier's own records rather than from activity, which is why the numbers there are smaller and duller than a usage dashboard's, and why they can be shown to somebody who checks.
Keep reading
- Where to put a human checkpoint on an agentFull autonomy is the wrong goal. Let an agent run through reversible steps and stop it before anything that cannot be undone: sending, deleting, charging.
- What replaces velocity when agents write codeStory points measured effort, and an agent can emit a hundred in an hour. What survives is verified throughput, first-pass rate and where the time went.
- What is AGENTS.md, and what belongs in it?AGENTS.md is a README for agents: the build commands, conventions and boundaries an agent needs. Over 60,000 repositories ship one. What to put in yours.