What is proof of work for AI agents?
Proof of work here has nothing to do with mining. It is the evidence a task carries to show the work was done, and the reason a status can be refused.
The phrase is borrowed, and the loan is unfortunate. Ask most people what proof of work means and they will describe a mining rig burning electricity to win a lottery. That is a different idea with the same name, and this article is about the other one: the ordinary English sense, where proof of work is the evidence that a piece of work was actually done.
That sense matters again because of who is doing the work. When a person moved a task to done, their claim came bundled with everything you already knew about them. An agent has none of that. It will report success in the same confident sentence whether it wrote the migration or wrote about writing the migration, and it will do it a hundred times a day.
Why a claim stopped being enough
A tracker is a record of claims. That was always true, and for twenty years it was fine, because the claims were made by people who had to sit in the same stand-up the next morning.
Two things break that. The first is volume: a team running agents produces more completed tasks per day than anyone will read carefully, so review degrades into sampling whether you designed it that way or not. The second is that the failure mode changed. A person who has not finished usually says so. A model that has not finished produces something that looks finished, because producing plausible-looking output is the thing it is best at.
So the question stops being "do I trust this worker" and becomes "what did this task actually produce, and can I look at it". That is proof of work.
What a proof of work is, concretely
In Cognibl it is a file, not a feeling. Specifically:
- A CSV, filled in against the template for the process the project runs.
- The documents it references, which is usually where the substance is.
- A coverage report, where the claim involves test cases.
The CSV is what makes the thing checkable rather than a note. A checker can ask whether the required columns are present and populated before a single word is read by a person. That is a cheap question to answer and it eliminates most of what would otherwise reach a human.
Harness, Graph and Loop each ship with a template, and a project running a custom process defines its own. This is worth being blunt about: the gate enforces the template you defined, so a template that asks for little will refuse little. The strictness is yours, not ours, and the report shows what was asked.
Versioned, never overwritten
Proof of work is versioned. A change mints the next version rather than editing the last, so v1 stays readable next to v3, and one version can hold several documents, because a real claim usually rests on more than one file.
That is not filing discipline for its own sake. The interesting question after something goes wrong is almost never "what does the evidence say now". It is "what did the evidence say when somebody approved this", and only an append-only record can answer it.
The gate is the point
Evidence nobody checks is paperwork. What makes proof of work load-bearing is that a status change which completes a task will not apply without it, and the refusal names what is missing rather than failing quietly.
Three layers do the checking, in order of cost:
- Deterministic checkers, for anything a machine can decide outright.
- An AI flow that reads the proof against the definition of done and writes up where they agree and where they do not.
- A human sample, at a rate set by the agent's score rather than by anyone's mood on the day.
An agent can attach proof and request the status change. It cannot decide the outcome, and there is no code path that special-cases our own agents either: a shortcut for them would be a bug, not a feature.
What it is not
It is not a copy of your work. Your systems stay the system of record, so code stays in your Git and documents stay in your Drive. What gets stored is a reference plus a content hash captured at verification time, which pins every report to the exact bytes it saw without a second copy existing anywhere to drift.
It is also not a substitute for writing down what you wanted. Proof of work is checked against a definition of done, and against a vague one it will confirm that something happened without telling you whether it was the right something. That is the harder half of the problem, and it is worth its own article.
Keep reading
- Where to put a human checkpoint on an agentFull autonomy is the wrong goal. Let an agent run through reversible steps and stop it before anything that cannot be undone: sending, deleting, charging.
- What replaces velocity when agents write codeStory points measured effort, and an agent can emit a hundred in an hour. What survives is verified throughput, first-pass rate and where the time went.
- What is AGENTS.md, and what belongs in it?AGENTS.md is a README for agents: the build commands, conventions and boundaries an agent needs. Over 60,000 repositories ship one. What to put in yours.