Where to put a human checkpoint on an agent
Full autonomy is the wrong goal. Let an agent run through reversible steps and stop it before anything that cannot be undone: sending, deleting, charging.
Put the checkpoint at the last reversible moment. Let an agent run without interruption through anything you can undo, then stop it before the step that cannot be taken back: sending a message, deleting data, publishing, merging, charging a card. This is what the field settled on during 2026, and it is a better rule than either "approve everything" or "let it run", both of which fail for the same reason.
Why is full autonomy the wrong target?
Because the goal was never autonomy. It was outcomes you do not have to check.
Teams that aim at full autonomy end up with a system nobody trusts, which means somebody reviews everything anyway, informally and without a record. Teams that approve every step build a queue that no human can service, and the approvals degrade into clicking yes. Both arrive at the same place: a control that exists on the diagram and not in practice.
The framing that works is selective autonomy. The agent is unsupervised where mistakes are cheap and reversible, and gated where they are not. That puts human attention where it changes an outcome instead of spreading it evenly over work that did not need it.
Which steps actually need a gate?
Reversibility first, then reach, then confidence.
Irreversibility is the primary test, and it is usually obvious. Reading a repository, drafting a change, running tests, proposing a branch: all recoverable, none worth a checkpoint. Merging to main, emailing a customer, deleting records, moving money, publishing anything: not recoverable by anyone, regardless of how sorry everybody is afterwards.
Reach is the second. An action that leaves your organisation deserves a gate even when it is technically reversible, because retraction is not the same as undo. A message sent to a customer has been read.
Confidence relative to stakes is the third and the most easily overdone. It is reasonable to gate on low model confidence for a high-stakes action. It is not reasonable to treat confidence as a substitute for the first two tests, because a confidently wrong agent is the normal failure mode rather than an unusual one.
What falls out is a short list that is stable across most products: send, delete, publish, merge, charge, grant access. If your gate list is much longer than that, you are gating steps, and the queue will eat the practice.
Is an approval the same as a done-gate?
No, and conflating them is the most common design error here.
An approval checkpoint asks a person to authorise an action before it happens. It is about permission, it runs at the moment of the act, and its question is "should this be allowed to proceed".
A done-gate asks whether the work satisfies what was asked. It runs after the work exists, its question is "does this evidence meet the definition of done", and it can be answered mechanically for most tasks because the criteria were written down in advance.
They fail differently, which is why you want both. Approvals catch the catastrophic and the irreversible. Done-gates catch the wrong-but-harmless, which is the far more common outcome and the one an approval queue is structurally bad at spotting. A person clicking approve on a merge is deciding whether merging is allowed, not reading the diff carefully at four in the afternoon.
How do you stop the queue becoming the bottleneck?
By measuring the queue, and by making the gate carry its own evidence.
Human wait share is the number that exposes this: the percentage of total cycle time spent waiting on a person rather than on work. It is usually the largest single component once agents are involved, and it responds to scheduling rather than to engineering, which makes it the cheapest thing available to fix.
Two design choices keep it small:
- Give the approver what they need to decide, in the request. An approval that requires opening three systems will be batched, delayed, and eventually rubber-stamped. The proposed change, what it affects, and what the checks said should be in front of them.
- Route by authority, not by availability. An approval sent to whoever is online is a formality. One sent to the person accountable for the thing is a decision.
And treat a rising approval queue as a signal about your gate list. If the queue grows steadily, you are probably gating reversible steps.
Does the regulation say anything about this?
It says the oversight has to work, which is a higher bar than having it.
The EU AI Act's Article 14 requires high-risk systems to be designed so people can effectively oversee them, with measures proportionate to the risk and the level of autonomy. The word doing the work there is effectively. A dashboard nobody owns is oversight in name. A checkpoint that pauses the action, and a halt that is one action rather than a deploy, are oversight in fact.
Whether that applies to you is a legal question with a real answer, and worth getting properly rather than from a blog. The design implication holds either way: build the stop before you need it, because a control you add during an incident is not a control.
What does this look like when it is working?
Quietly, mostly. The agent works for long stretches without anyone watching, because the steps it is taking cannot hurt anybody. Attention arrives at a small number of moments that are genuinely decisions. And the number of approvals per unit of work falls over time as the reversible envelope gets better understood, rather than rising as trust erodes.
In Cognibl the reversible envelope is structural rather than configured. An agent's write-type tool calls land in a staging queue instead of in your systems, so the default state of an agent's work is proposed rather than applied, and the commit step is what pushes it through or discards it. Halting is one action at the job, agent or customer level, and never a deployment.
Keep reading
- What replaces velocity when agents write codeStory points measured effort, and an agent can emit a hundred in an hour. What survives is verified throughput, first-pass rate and where the time went.
- What is AGENTS.md, and what belongs in it?AGENTS.md is a README for agents: the build commands, conventions and boundaries an agent needs. Over 60,000 repositories ship one. What to put in yours.
- How much agent work should a human review?Reviewing everything stops being possible long before anyone decides to stop. Sampling deliberately, at a rate set by evidence, beats sampling by accident.