Skip to content
All writing

Writing a definition of done an AI agent can pass

Aarav Pundir · FounderDefinition of done4 min

A definition of done written for a person is a reminder. Written for an agent it is a specification, and the difference shows up in your first-pass rate.

Every tracker has had a definition-of-done field for a decade, and on most teams it holds something like "works and is tested". That was never a specification. It was a reminder, addressed to a colleague who already knew what the task was for, and it worked because the reader filled in everything it left out.

An agent fills in the gaps too. It just fills them in with something plausible rather than something correct, and it does so without the pause a person would have taken to ask.

The reminder and the specification

Here is the distinction that matters. A reminder is read by someone who shares your context; it needs only to jog them. A specification is read by someone who has none; it has to survive being taken literally.

"Handle the edge cases" is a reminder. Taken literally it is unbounded, so it gets satisfied by whatever the reader happened to think of first. "Rejects an empty file, a file over 10MB, and a row whose date does not parse, and reports which row failed" is a specification. It can be wrong, which is exactly what makes it useful.

That is the test, and it is a good one to apply to anything you have written: could this be marked failed? If no honest reading of the sentence could ever fail, it is not doing any work.

Write the check, not the intent

The practical move is to write down the thing that would be observed, rather than the state of the world you are hoping for.

Intent sounds like: the import should be reliable. Robust error handling. Good test coverage. Nothing here is checkable, so nothing here is refusable.

A check sounds like: a malformed row leaves the other rows imported and appears in a rejects file with its line number. The test suite covers the three rejection paths. A run over the sample file completes without manual intervention.

The second version is longer, and that is the actual cost. It gets paid once, up front, by the person who understands the task best, instead of repeatedly by whoever is reviewing output that nearly works.

Say which systems may be touched

A definition of done written for an agent has a second job a human-facing one never needed: bounding the blast radius. A person knows not to refactor the billing module while fixing a typo. An agent knows nothing of the sort unless the task says so.

So name the surface. Which repository, which branch, which files or directories are in scope, and what must not be touched. This is not paranoia, it is the same instinct that makes you scope a pull request, written down for a worker who cannot infer it.

Watch two numbers together

The reason to care about any of this is measurable, and it takes two metrics rather than one.

First-pass verification rate is the share of agent tasks that pass on the first attempt. It is the specification-quality metric: a well-written definition of done passes first time, and a vague one is discovered here rather than in production.

Spec time is the part of cycle time from creation to first run, including every revision of the definition of done.

Neither means much alone. Long spec time is not a fault by itself; some tasks are genuinely hard to specify and the thinking has to happen somewhere. Long spec time beside a low first-pass rate is the signal worth acting on: it means the definitions are being written slowly and still not landing, which is a sign the work is being described rather than specified.

The reverse pattern is worth noticing too. A very high first-pass rate with almost no spec time usually means the definitions are too loose to fail, and the gate is confirming that something happened rather than that the right thing did.

A shape that works

Four parts, in this order:

  1. The observable outcome. What is true afterwards that is not true now, stated so it could be checked by someone who was not in the conversation.
  2. The checks. What specifically will be run or looked at, including the cases that should fail.
  3. The surface. Which systems and files are in scope, and which are not.
  4. The evidence. What the task should produce so the outcome can be confirmed later, which is the proof of work.

It is not a long document. Most tasks fit in a short paragraph and a handful of bullets. What it must not be is a sentence that no outcome could contradict.

Start where it hurts

You do not need to rewrite your backlog. Take the last five tasks that came back wrong and read what their definitions of done actually said. The pattern is usually immediate, and it is usually the same: the sentence described the work rather than the result, so anything that resembled the work satisfied it.