Skip to content
All writing

Why AI agent failures are context failures

Aarav Pundir · FounderContext4 min

Agents rarely fail because the model was not capable. They fail because nobody told them what done meant. The 2026 data now says this out loud.

When an agent produces the wrong thing, the cause is almost never that the model could not do the work. It is that the model was working from an incomplete picture of what the work was: the constraint nobody wrote down, the convention everyone on the team knows, the definition of done that said "fix the bug" and stopped there. The industry spent two years upgrading models and has now largely concluded the bottleneck was somewhere else.

What does the evidence actually say?

Redis's 2026 context engineering survey put the question directly to practitioners and found 73% saying agents fail more often from broken context than from broken models, with 83% saying fresh context matters more for enterprise reliability than adding model parameters.

The more interesting half of that finding is the gap that follows it. Belief has moved; practice has not. The same survey describes most organisations still assembling context by hand, with no shared infrastructure, no freshness guarantees and little governance. People have identified the problem and are mostly still solving it with copy and paste.

Why does a capable model still get it wrong?

Because it resolves ambiguity instead of surfacing it.

Give an underspecified task to a competent colleague and you get a question, or at worst a guess they mention. Give the same task to a model and you get a confident, complete-looking answer built on an assumption it never announced. That is not a defect in the model. Producing plausible continuations is the thing it does, and an ambiguous instruction has plausible continuations.

So the failure mode inverts. A person who has not understood usually looks like they have not understood. An agent that has not understood looks exactly like an agent that has, which is why review by eye stops working at volume and why verification has to be structural rather than attentive.

Is this the same as prompt engineering?

No, and the distinction is why the term changed.

Prompt engineering was about phrasing: the wording, the examples, the role instruction. It treats the model as the variable. Context engineering treats the information as the variable, and asks a different question: what does this worker need to know, where does it come from, how does it stay current, and who decided it was allowed to see it.

That reframing is what makes it an infrastructure problem rather than a writing problem. Phrasing does not have a freshness guarantee. A retrieval path does. A scope does. A definition of done written before the work does, and it can be version-controlled, which a clever prompt in somebody's history cannot.

Where does the missing context usually live?

Four places, in roughly descending order of how often they bite:

  • In the acceptance criteria that were never written. The task says what to build and not how anyone will decide it worked.
  • In the team's head. Conventions, the reason a thing is done the odd way, the module you do not touch before the migration lands.
  • In an adjacent system the agent cannot see. The ticket references a design, the design lives elsewhere, the agent has neither.
  • In the run itself, then lost. Long sessions degrade: information that was present early stops being usable late, and a multi-step task quietly drifts from what it was originally asked to do.

The first is the largest and the most fixable, which is fortunate, because it is also the one nobody wants to do. Writing a checkable done-test is slower than writing a task title, and it is the single highest-leverage thing available.

Does more context fix it?

Not reliably, and this is where teams overcorrect.

The instinct after a context failure is to send more: the whole repository, the full history, every related document. It helps less than expected and sometimes hurts, because relevance falls as volume rises and long sessions degrade rather than improve. Stuffing is not engineering.

What works is narrower and duller. Send what the task needs, from a source that is current, with the acceptance criteria stated in terms something can check. The goal is not a bigger picture. It is a picture with the decisive part in it.

What does this change about how you run agents?

It moves the effort earlier and it makes the effort visible.

DORA's 2025 report found AI adoption improving throughput while delivery stability declined, and named the phenomenon teams keep rediscovering privately: time saved writing gets re-spent auditing. If the specification was thin, the audit is where you pay for it, at the least convenient moment and usually with a person's attention.

Which is why the useful measurements are the ones that watch the front of the process rather than the end. First-pass verification rate is a specification quality metric. Cycle time split by phase shows whether the spec phase is investment or rework. Both are boring numbers that make context debt legible before it converts into an incident.

Cognibl is built on the same premise: a task carries its definition of done, its scope and its evidence together, so the context that decides whether the work counts travels with the work instead of living in somebody's memory of the stand-up.