Skip to content
All writing

Agent governance vs observability vs verification

Aarav Pundir · FounderCategory4 min

Three tools that sound alike and answer different questions: what happened, what is allowed, and whether the work is actually acceptable. Only one gates done.

Observability tells you what an agent did. Governance decides what it is allowed to do. Verification decides whether what it produced is acceptable. Three questions, three different moments in time, and buying one does not get you the others, which is the mistake this article exists to prevent.

What does each one actually do?

Observability is retrospective. It captures telemetry across an agent's reasoning and execution: the prompt, the tool calls, the responses, the tokens, the latency. Its output is a trace you read after the fact. Current definitions of the field describe it as monitoring, tracing, evaluation and governance gathered into one telemetry practice.

Governance is concurrent. It decides, at the moment of the call, whether this agent may use this tool with this scope. Its output is an allow or a refusal. Policy that only exists in a document is not governance; policy enforced at execution time is.

Verification is terminal. It runs after the work is produced and before it counts, and it answers one question: does this output satisfy the definition of done it was contracted against? Its output is a verdict, and the verdict is what a status change depends on.

The timing is the whole distinction. Observability happens after and cannot refuse. Governance happens during and cannot judge quality. Verification happens at the boundary and does both badly if you ask it to do either of the others.

Why is verification the one that is usually missing?

Because it is the only one that has to say no to your own team.

Observability is easy to adopt: it adds information and threatens nothing. It makes engineers faster at debugging and gives leadership a dashboard. Governance is harder but familiar, because it looks like access control, which every organisation already has opinions about.

Verification is different in kind. It is a gate that refuses work, which means it will refuse work somebody wanted shipped, on a day when shipping mattered. Every organisation that installs one discovers this in the first fortnight. It is also the one whose absence is invisible: nothing breaks, the traces still flow, the policies still hold, and completed tasks quietly accumulate that nobody checked.

The measurement data suggests this is close to universal. Arcade.dev's 2026 survey found 52% of gen-AI organisations running agents in production and 31% with any measurement framework for them. The gap is not a tooling gap so much as a gap in what people thought they had bought.

Do the three overlap?

At the edges, and the overlaps are where the confusion lives.

Observability platforms increasingly ship evaluation features, which look like verification. The difference is authority: an evaluation produces a score you can look at, and verification produces a verdict something enforces. A score with no gate behind it is observability with a rubric.

Governance platforms increasingly ship audit trails, which look like observability. That one is largely real, because a gateway that enforces policy is sitting in the ideal position to record what passed through it.

The overlap that does not exist is governance and verification. Knowing an agent was allowed to open a pull request tells you nothing about whether the pull request is any good. Deny-by-default and done-gating are unrelated mechanisms solving unrelated failures, and a vendor collapsing them is worth a question.

How do the three appear on a single task?

Take one ordinary unit of work: an agent is asked to fix a bug.

  1. Governance, at the start. The agent gets a scoped key. It may read the repository and open a branch. It may not merge, and it may not touch billing. Anything outside that is refused at the call.
  2. Observability, throughout. Every tool call it makes is recorded with its arguments, its result and its cost. If the run behaves oddly, this is what somebody reads.
  3. Verification, at the end. The agent says it is done and attaches its evidence. The gate checks that evidence against the definition of done written before the work started, and returns pass or fail. Only a pass moves the status.

Remove step one and the agent can do anything it is clever enough to reach. Remove step two and you cannot investigate what it did. Remove step three and every task is done the moment the worker says so, which is the state most agent programmes are actually in.

Which should you buy first?

Buy in the order of what the failure would cost you.

If an agent has write access to something that matters, governance is first, because that failure is irreversible and the others are not. If the work is low-stakes but you cannot explain what happened when it goes wrong, observability is first, because you are debugging blind.

Verification becomes first the moment agent output is entering a system of record faster than a person can read it. That is the point where sampling stops being a policy and becomes an accident, and where "we review everything" quietly becomes false without anybody deciding it.

Cognibl is a verification-first product with a gateway attached: the done-gate is the thing we are, the scoped key and the trace are what make its verdicts mean anything. We are deliberately not an observability platform, and if what you need is deep tracing of model behaviour, the specialists are better at it than we are.