What is agent washing?
Agent washing is what happens when a product gets relabelled rather than rebuilt. A chatbot, an assistant, or a scripted automation acquires agentic marketing language without acquiring the ability to decide its own steps. Gartner named the practice in June 2025.
The lineage is obvious. Greenwashing described environmental claims that outran the practice; AI washing described products that claimed machine learning and delivered rules. Agent washing is the same move applied to a newer word, and the stakes are higher because an agent is bought for autonomy that a chatbot cannot supply at any price.
Gartner's June 2025 assessment put a number on it that has been quoted ever since: of the thousands of vendors claiming agentic AI, only around 130 were judged to offer genuine agentic capability. Whatever you make of the precision of that figure, the direction is not seriously disputed by anyone selling into this market.
The uncomfortable part: much of it is not deception
It is tempting to read agent washing as vendors lying, and some of it is. The more useful reading is that there is no standards body for what qualifies as an agent, no certification, and no agreed threshold. In that vacuum, every vendor sets the bar where their product already stands, and most of them believe their own framing.
That changes the remedy. If the problem were dishonesty, you would solve it by finding honest vendors. Because the problem is definitional, you cannot rely on the label at all, from anyone, including from vendors acting in complete good faith. You need a test you apply yourself. The test that does most of the work is on the AI agent page: who decided the order of the steps?
Why does agent washing happen?
Four forces, only one of which is dishonesty. The word has no agreed boundary, buyers ask for agents by name, budget follows the label rather than the capability, and a genuine agent is considerably harder to build than a chatbot with a few functions attached.
Understanding the mechanics matters because it tells you which signals are meaningful. A vendor responding rationally to market incentives behaves differently from one attempting fraud, and both produce a deck that says "agentic."
- No boundary to cross. Nobody can be caught out on a definition that does not exist. A product that plans one step and calls one tool sits somewhere on the spectrum of agency, and the vendor is not technically wrong to say so.
- Buyers ask by name. When RFPs specify agentic AI, a vendor whose product genuinely suits the problem but is not agentic has a choice between losing the deal and adjusting the language. Demand pulls the label.
- Budget attaches to the word. Agentic AI unlocks funding that "workflow automation improvements" does not, at both vendor and buyer. That incentive operates on your own teams too, which the section below covers.
- The real thing is hard. An agent that decides its own path needs planning, memory, tool orchestration, a stopping condition, guardrails, evaluation, and a governance story. Adding an agentic label to a working chatbot takes a marketing sprint.
The picture above is the whole problem in one image. The same two words are applied across a capability range that spans an order of magnitude, which means the label carries no information. It is not that the label is often wrong. It is that it cannot be right or wrong, because it excludes nothing.
How do you detect agent washing?
Ask five questions that a demonstration cannot answer. Each targets a capability that has to be built rather than claimed, and a washed product tends to fail on the same two: the stopping condition and the trace of a failed run.
Demonstrations are rehearsed and prove almost nothing about autonomy, because a scripted path and a chosen path look identical when the script was written for that input. These questions are designed to be answerable only with evidence.
- 01
Who decided the order of the steps?If a developer or analyst wrote the sequence, it is automation with AI inside it, whatever the label. What you are listening for: a straight answer. Deflection to "it depends on configuration" usually means a person configured it.
- 02
Show me a run where it did something you did not anticipateGenuine agency produces surprises, including good ones. What you are listening for: a specific example with a trace. A vendor who has never seen their product do anything unexpected is describing a workflow engine.
- 03
What decides when it stops?The stopping condition is the component most often absent, and the one no demonstration reveals, because demonstrations end when the presenter stops them. What you are listening for: named limits on steps, time, and spend, and a definition of done.
- 04
Show me the trace of a run that failedThis is the single most reliable tell. A real agent platform has per-step traces because it cannot be operated without them. What you are listening for: tool calls, arguments, and the decision points. Screenshots of a dashboard are not a trace.
- 05
What happens if I remove one of its tools?An agent adapts and finds another route or reports that it cannot proceed. A scripted product breaks at that step. What you are listening for: a willingness to try it in front of you.
Questions three and four do most of the work. They are the two capabilities that cannot be retrofitted into a demo, because both require the product to have been built for production rather than for a pitch. Ask them early. If the answers are vague, the remaining questions will be too, and you have saved yourself an evaluation cycle.
What does agent washing actually cost?
More than the licence fee. A washed purchase underdelivers, the organization concludes that agentic AI does not work, and the next proposal is harder to fund. The damage is to the category inside your own company, and it outlasts the contract.
The direct cost is easy to see and rarely the largest one. Four consequences follow a washed purchase, in roughly this order.
- You pay an agentic price for automation. The immediate loss, and the most recoverable. Contracts end.
- The pilot underdelivers and the diagnosis is wrong. When the results disappoint, the failure is attributed to agentic AI rather than to the product that was not agentic. The wrong lesson gets learned and written down.
- The category is poisoned internally. The next team proposing genuine agentic work now argues against an internal precedent. This is the expensive part, it does not appear in any budget line, and it can set a programme back a year.
- The skepticism becomes self-reinforcing. Washed products fail, buyers lose confidence in the category, and legitimate vendors face heightened suspicion, which raises the cost of every subsequent evaluation for everyone.
There is a variant worth naming separately, because it is not washing and it produces the same outcome. A vendor can describe their product accurately and still leave you disappointed if you assumed more autonomy than the description implied. Honest description and accurate expectation are two different things, and the five questions above close the gap in both cases.
Does agent washing happen inside enterprises too?
Yes, and it is discussed far less than the vendor version. Internal teams relabel existing automation as agentic to secure funding, because budget follows the word. The organization then counts projects as agentic that were never agentic, and draws conclusions from the wrong denominator.
The incentive is identical to the vendor's and the mechanism is the same. A team with a working rules engine and a funding round to survive describes it in the language that gets funded. Nobody involved feels dishonest, and in many cases the relabelling is a fair description of an ambition rather than of a system.
Two consequences follow, and both are worse than they first appear.
- Your agent inventory becomes fiction. If a third of what is registered as an agent is scripted automation, then the count, the risk assessment, and the governance requirements attached to it are all wrong in ways nobody can see. This is a different failure from agent sprawl and it compounds with it: you cannot say how many agents you run, and some of what you have counted is not an agent.
- Your cancellation rate misleads you. When relabelled automation is cancelled because it did not deliver agentic outcomes, that cancellation is recorded against agentic AI. Some proportion of every reported failure statistic, internal and industry-wide, is measuring agent washing being discovered rather than agentic AI failing.
How should you read agent washing statistics?
Check the date before you quote the number. The figures in circulation come from a single Gartner release in June 2025, and much of the 2026 coverage repeats them without the date, which makes a year-old prediction read as fresh research.
This is worth doing carefully on a page about overclaiming, because the statistics used to warn about hype are themselves being recirculated in a way that overclaims. Here is the provenance of every figure you are likely to encounter.
| The figure | What it actually is |
|---|---|
| Over 40% of agentic AI projects cancelled by end of 2027 | A Gartner prediction, published 25 June 2025. Not an observed outcome, and not new. Causes given: escalating costs, unclear business value, inadequate risk controls |
| Only ~130 of thousands of vendors are real | A Gartner estimate from the same June 2025 release. An analyst judgement rather than a measured count, and the vendor population has moved since |
| 19% significantly invested, 42% cautious | A Gartner poll of 3,412 webinar attendees, January 2025. A self-selecting audience, so treat it as directional |
| 15% of day-to-day work decisions autonomous by 2028 | From the same release, and rarely quoted alongside the 40% figure because it points the other way |
| 33% of enterprise software will include agentic AI by 2028 | Also from June 2025, up from under 1% in 2024. Note that this measures inclusion of a feature, not agentic outcomes |
Two things follow. The same source that predicts the cancellations also predicts substantial adoption, and coverage that quotes only the pessimistic half is selecting. Both are in one press release, and read together they describe an ordinary adoption curve rather than a collapse.
And a statistic quoted without its date is doing rhetorical work. When a mid-2026 article presents a June 2025 prediction as a current finding, that is the same move as agent washing performed on a number instead of a product: an existing thing relabelled to seem newer than it is. The habit that protects you is small. Ask when the figure was published, what it measured, and whether it was a prediction or an observation, before it reaches a board slide with your name on it.