Skip to content
All writing

How to tell a real AI agent from agent washing

Aarav Pundir · FounderAgent washing3 min

Agent washing means calling a chatbot or a script an AI agent. Gartner expects it to help cancel over 40% of agentic projects by 2027. This is how to spot it.

Agent washing means taking an existing product and calling it an "AI agent" without giving it the autonomy the word now implies. The product might be a chatbot, a robotic process automation script, or a wrapper around one model call. The term comes from Gartner, and so does a stark number. Gartner expects over 40% of agentic AI projects to be canceled by the end of 2027. It also estimates that of the thousands of vendors selling "agentic" today, only around 130 are the real thing. The rest are agent washing.

What is agent washing?

It is a change in marketing with no change in the technology. The product is described in the language of autonomous work: deciding, planning, acting across systems. What it does is answer a prompt or run a fixed script. The new label adds no new capability.

Gartner's figure came from a poll of over 3,400 organisations investing in agentic AI. It gives three specific reasons so many projects will be scrapped: escalating cost, unclear business value, and inadequate risk controls. None of the three is a problem with the model. Each one comes from buying a capability that was never there to measure.

Why is it suddenly everywhere?

Because "agent" sells, and almost nothing checks the claim before a buyer signs the purchase order.

There is no standard for what an agent has to be able to finish, so the word is not tied to any test. A buyer comparing ten "agentic" products is comparing ten slide decks instead of ten verified outcomes. With no test, vendors have every reason to relabel. The demo goes well, and the invoice is the same either way.

How do you tell a real agent from a relabelled one?

Stop asking what it can start. Ask what it can finish, and what it leaves behind when it does.

A relabelled chatbot produces output: text, a suggestion, a draft. A real agent completes a unit of work against a stated definition of done and leaves evidence that a third party can check without having watched it work. The difference is a done-test you can check, and a record of the run against it. If the only artefact is the output, and nothing says whether the output was correct, you are looking at agent washing. That holds however smooth the demo was.

Why do these projects get canceled?

All three of Gartner's reasons point to the same gap.

Escalating cost with unclear business value happens when nobody can say, in numbers, what the agent finished that a person did not have to redo. Inadequate risk controls happen when an agent has access to a live system and no gate stands between its output and the customer's data. Both come from a gap in measurement and verification. A stronger model would not fix either one. A project that cannot show what its agents completed cannot defend its budget when the review comes.

What protects a buyer?

Accept agent work on proof, not on the strength of a demo.

That means three things a relabelled product cannot fake. Work is verified before it is billed, so an invoice line points at a passing report rather than at raw usage. Access is denied by default, so an agent reaches only the tools its contract named. And the log is written by the system the agent calls through, not self-reported by the agent. The record of what happened does not depend on the vendor's word. A product built this way gains nothing from relabelling, because every claim it makes comes with its evidence.

Does this mean agents are overhyped?

No. It helps to judge the category separately from the vendors selling into it.

The 40% of projects that get canceled and the roughly 130 real builders are two views of the same market. Agents that finish real work exist and are shipping. The cancellations show that most failures are governance failures rather than model failures. The work was never specified in a way anything could check, so nobody could tell finished work from work that only looked finished. The field keeps reaching the same finding from other directions: agent failures are usually context failures. That is a problem you can fix without walking away from the technology.

Cognibl is built so that agent washing cannot pass as real work. A task carries its definition of done, its scope and its evidence together. No status reaches "done" without a proof of work behind it. You do not have to take the agent's word, or ours. You read the report. See what a verified completion looks like.