All glossary terms
A Market reality Procurement

Agent washing

Agent washing is relabelling a product as agentic AI without rebuilding it. The word "agent" has no agreed boundary, so vendors set the threshold wherever their product happens to sit. Which means you need a test, not a definition.

Definition

Agent washing is the rebranding of existing products as agentic AI without substantial agentic capability: chatbots, assistants, and scripted automation relabelled rather than rebuilt. Gartner named the practice in June 2025, alongside its estimate that only around 130 of the thousands of vendors claiming agentic AI offered genuine capability.

What is agent washing?

Agent washing is what happens when a product gets relabelled rather than rebuilt. A chatbot, an assistant, or a scripted automation acquires agentic marketing language without acquiring the ability to decide its own steps. Gartner named the practice in June 2025.

The lineage is obvious. Greenwashing described environmental claims that outran the practice; AI washing described products that claimed machine learning and delivered rules. Agent washing is the same move applied to a newer word, and the stakes are higher because an agent is bought for autonomy that a chatbot cannot supply at any price.

Gartner's June 2025 assessment put a number on it that has been quoted ever since: of the thousands of vendors claiming agentic AI, only around 130 were judged to offer genuine agentic capability. Whatever you make of the precision of that figure, the direction is not seriously disputed by anyone selling into this market.

The uncomfortable part: much of it is not deception

It is tempting to read agent washing as vendors lying, and some of it is. The more useful reading is that there is no standards body for what qualifies as an agent, no certification, and no agreed threshold. In that vacuum, every vendor sets the bar where their product already stands, and most of them believe their own framing.

That changes the remedy. If the problem were dishonesty, you would solve it by finding honest vendors. Because the problem is definitional, you cannot rely on the label at all, from anyone, including from vendors acting in complete good faith. You need a test you apply yourself. The test that does most of the work is on the AI agent page: who decided the order of the steps?

Why does agent washing happen?

Four forces, only one of which is dishonesty. The word has no agreed boundary, buyers ask for agents by name, budget follows the label rather than the capability, and a genuine agent is considerably harder to build than a chatbot with a few functions attached.

Understanding the mechanics matters because it tells you which signals are meaningful. A vendor responding rationally to market incentives behaves differently from one attempting fraud, and both produce a deck that says "agentic."

  • No boundary to cross. Nobody can be caught out on a definition that does not exist. A product that plans one step and calls one tool sits somewhere on the spectrum of agency, and the vendor is not technically wrong to say so.
  • Buyers ask by name. When RFPs specify agentic AI, a vendor whose product genuinely suits the problem but is not agentic has a choice between losing the deal and adjusting the language. Demand pulls the label.
  • Budget attaches to the word. Agentic AI unlocks funding that "workflow automation improvements" does not, at both vendor and buyer. That incentive operates on your own teams too, which the section below covers.
  • The real thing is hard. An agent that decides its own path needs planning, memory, tool orchestration, a stopping condition, guardrails, evaluation, and a governance story. Adding an agentic label to a working chatbot takes a marketing sprint.
A capability spectrum running left to right: scripted chatbot, chatbot with function calling, assistant with a fixed workflow, bounded routing, and finally a system that plans its own path. A single uniform banner reading agentic AI is stretched across the entire range, showing that one label covers wildly different capability. A note reads that the label carries no information because it spans everything.

The picture above is the whole problem in one image. The same two words are applied across a capability range that spans an order of magnitude, which means the label carries no information. It is not that the label is often wrong. It is that it cannot be right or wrong, because it excludes nothing.

How do you detect agent washing?

Ask five questions that a demonstration cannot answer. Each targets a capability that has to be built rather than claimed, and a washed product tends to fail on the same two: the stopping condition and the trace of a failed run.

Demonstrations are rehearsed and prove almost nothing about autonomy, because a scripted path and a chosen path look identical when the script was written for that input. These questions are designed to be answerable only with evidence.

Five questions arranged as a scorecard with what a genuine agent can show and what a washed product typically cannot. Who decided the order of the steps. Show me a run where it did something you did not anticipate. What decides when it stops. Show me the trace of a run that failed. What happens if I remove one of its tools. The stopping condition and the failed trace are marked as the two most reliable tells.
  1. 01
    Who decided the order of the steps?If a developer or analyst wrote the sequence, it is automation with AI inside it, whatever the label. What you are listening for: a straight answer. Deflection to "it depends on configuration" usually means a person configured it.
  2. 02
    Show me a run where it did something you did not anticipateGenuine agency produces surprises, including good ones. What you are listening for: a specific example with a trace. A vendor who has never seen their product do anything unexpected is describing a workflow engine.
  3. 03
    What decides when it stops?The stopping condition is the component most often absent, and the one no demonstration reveals, because demonstrations end when the presenter stops them. What you are listening for: named limits on steps, time, and spend, and a definition of done.
  4. 04
    Show me the trace of a run that failedThis is the single most reliable tell. A real agent platform has per-step traces because it cannot be operated without them. What you are listening for: tool calls, arguments, and the decision points. Screenshots of a dashboard are not a trace.
  5. 05
    What happens if I remove one of its tools?An agent adapts and finds another route or reports that it cannot proceed. A scripted product breaks at that step. What you are listening for: a willingness to try it in front of you.

Questions three and four do most of the work. They are the two capabilities that cannot be retrofitted into a demo, because both require the product to have been built for production rather than for a pitch. Ask them early. If the answers are vague, the remaining questions will be too, and you have saved yourself an evaluation cycle.

What does agent washing actually cost?

More than the licence fee. A washed purchase underdelivers, the organization concludes that agentic AI does not work, and the next proposal is harder to fund. The damage is to the category inside your own company, and it outlasts the contract.

The direct cost is easy to see and rarely the largest one. Four consequences follow a washed purchase, in roughly this order.

  • You pay an agentic price for automation. The immediate loss, and the most recoverable. Contracts end.
  • The pilot underdelivers and the diagnosis is wrong. When the results disappoint, the failure is attributed to agentic AI rather than to the product that was not agentic. The wrong lesson gets learned and written down.
  • The category is poisoned internally. The next team proposing genuine agentic work now argues against an internal precedent. This is the expensive part, it does not appear in any budget line, and it can set a programme back a year.
  • The skepticism becomes self-reinforcing. Washed products fail, buyers lose confidence in the category, and legitimate vendors face heightened suspicion, which raises the cost of every subsequent evaluation for everyone.

There is a variant worth naming separately, because it is not washing and it produces the same outcome. A vendor can describe their product accurately and still leave you disappointed if you assumed more autonomy than the description implied. Honest description and accurate expectation are two different things, and the five questions above close the gap in both cases.

Does agent washing happen inside enterprises too?

Yes, and it is discussed far less than the vendor version. Internal teams relabel existing automation as agentic to secure funding, because budget follows the word. The organization then counts projects as agentic that were never agentic, and draws conclusions from the wrong denominator.

The incentive is identical to the vendor's and the mechanism is the same. A team with a working rules engine and a funding round to survive describes it in the language that gets funded. Nobody involved feels dishonest, and in many cases the relabelling is a fair description of an ambition rather than of a system.

Two consequences follow, and both are worse than they first appear.

  • Your agent inventory becomes fiction. If a third of what is registered as an agent is scripted automation, then the count, the risk assessment, and the governance requirements attached to it are all wrong in ways nobody can see. This is a different failure from agent sprawl and it compounds with it: you cannot say how many agents you run, and some of what you have counted is not an agent.
  • Your cancellation rate misleads you. When relabelled automation is cancelled because it did not deliver agentic outcomes, that cancellation is recorded against agentic AI. Some proportion of every reported failure statistic, internal and industry-wide, is measuring agent washing being discovered rather than agentic AI failing.

How should you read agent washing statistics?

Check the date before you quote the number. The figures in circulation come from a single Gartner release in June 2025, and much of the 2026 coverage repeats them without the date, which makes a year-old prediction read as fresh research.

This is worth doing carefully on a page about overclaiming, because the statistics used to warn about hype are themselves being recirculated in a way that overclaims. Here is the provenance of every figure you are likely to encounter.

The commonly quoted agent washing and agentic AI statistics, with their source and original publication date
The figure What it actually is
Over 40% of agentic AI projects cancelled by end of 2027 A Gartner prediction, published 25 June 2025. Not an observed outcome, and not new. Causes given: escalating costs, unclear business value, inadequate risk controls
Only ~130 of thousands of vendors are real A Gartner estimate from the same June 2025 release. An analyst judgement rather than a measured count, and the vendor population has moved since
19% significantly invested, 42% cautious A Gartner poll of 3,412 webinar attendees, January 2025. A self-selecting audience, so treat it as directional
15% of day-to-day work decisions autonomous by 2028 From the same release, and rarely quoted alongside the 40% figure because it points the other way
33% of enterprise software will include agentic AI by 2028 Also from June 2025, up from under 1% in 2024. Note that this measures inclusion of a feature, not agentic outcomes

Two things follow. The same source that predicts the cancellations also predicts substantial adoption, and coverage that quotes only the pessimistic half is selecting. Both are in one press release, and read together they describe an ordinary adoption curve rather than a collapse.

And a statistic quoted without its date is doing rhetorical work. When a mid-2026 article presents a June 2025 prediction as a current finding, that is the same move as agent washing performed on a number instead of a product: an existing thing relabelled to seem newer than it is. The habit that protects you is small. Ask when the figure was published, what it measured, and whether it was a prediction or an observation, before it reaches a board slide with your name on it.

Frequently asked questions about agent washing

If your question is not here, our team will answer it directly.


Talk to a Specialist →
What is agent washing in simple terms?

You are sold an agent and you receive a chatbot with a new price tag. A product that already existed, typically an assistant, a chatbot, or a scripted automation, gets described in agentic language without gaining the ability to decide its own steps. The name follows greenwashing and AI washing, and the stakes are higher here because autonomy is precisely what the buyer is paying extra for, and it is the one thing a relabelled product cannot supply.

Who coined the term agent washing?

Gartner, in a press release dated 25 June 2025, which described it as relabelling AI assistants, robotic process automation and chatbots as agentic without adding the underlying capability. That release also carried the two numbers now quoted constantly: a judgement that roughly 130 vendors out of the thousands marketing the category were genuinely delivering it, and a forecast that more than 40% of agentic AI projects would be cancelled by the end of 2027. Both are usually repeated with the date stripped off.

How can you tell if a product is really an AI agent?

Ask five questions a demonstration cannot answer. Who decided the order of the steps? Show me a run where it did something you did not anticipate. What decides when it stops? Show me the trace of a run that failed. What happens if I remove one of its tools? The third and fourth do most of the work, because a stopping condition and per-step traces both have to be built for production and cannot be retrofitted into a pitch.

Is agent washing always deliberate?

No, and assuming it is leads you to the wrong remedy. There is no standards body for what qualifies as an agent, no certification, and no agreed threshold, so vendors set the bar where their product already stands and most believe their own framing. If the problem were dishonesty you would solve it by finding honest vendors. Because it is definitional, you cannot rely on the label from anyone, including from vendors acting in complete good faith, which is why you need your own test.

Why does agent washing matter if the product still works?

Because the damage outlasts the contract. You pay an agentic price for automation, which is recoverable. Then the pilot underdelivers and the failure gets attributed to agentic AI rather than to a product that was never agentic, so the wrong lesson is learned and written down. The next team proposing genuine agentic work now argues against an internal precedent. That third cost appears in no budget line and can set a programme back a year.

Does agent washing happen inside companies, not just from vendors?

Yes, and it is discussed far less. Internal teams relabel existing automation as agentic because budget follows the word, and the incentive is identical to the vendor's. Two consequences follow. Your agent inventory becomes fiction, so the count, the risk assessment, and the governance attached to it are wrong in ways nobody can see. And when relabelled automation is cancelled for not delivering agentic outcomes, the cancellation is recorded against agentic AI.

Is the Gartner 40% cancellation statistic still accurate?

It was a prediction, not an observation, and it was published on 25 June 2025. Much of the coverage circulating in 2026 repeats it without the date, which makes a year-old forecast read as fresh research. It has not been superseded, and it also has not been confirmed, since the window it describes runs to the end of 2027. Before quoting it, establish when it was published, what it measured, and whether it was a forecast or a finding.

Are only 130 agentic AI vendors real?

That was Gartner's estimate in June 2025, and it is an analyst judgement rather than a measured count. Treat it as a statement about the ratio rather than about the exact number, because the vendor population has changed since and no register exists against which such a figure could be verified. The useful reading is the proportion it implies: the large majority of products marketed as agentic were assessed as not being agentic, which is a claim few people selling into this market seriously dispute.

Is automation with AI in it a bad thing?

Not at all, and for many processes it is the better engineering choice. A fixed path with a model handling one interpretive step is cheaper, faster, easier to test, and can be approved once rather than supervised continuously. The problem with agent washing is not that the underlying product is poor. It is that it is sold and priced as something it is not, and then governed as something it is not, which produces disappointment on both sides of a transaction that might otherwise have suited everyone.

What is the difference between agent washing and AI washing?

The same behaviour applied to a different claim. AI washing described products that claimed machine learning while delivering rules-based logic, and it has attracted regulatory attention in some markets on the basis that it constitutes misleading investor or customer disclosure. Agent washing is the successor: the product may genuinely use AI and still not be agentic, because the claim at stake is autonomy rather than intelligence. A washed agent can contain real AI and still not decide anything.

How do you avoid buying an agent-washed product?

Stop evaluating against the label and evaluate against the autonomy you actually need. Decide first whether your process requires a system that chooses its own path, because many do not. Then apply the five questions, weighting the stopping condition and the failed-run trace. Finally, write your requirement in capability terms rather than category terms: a specification asking for agentic AI invites a relabelled response, while one asking for a system that selects its own tool sequence and produces per-step traces does not.

Will agent washing decline over time?

Probably, and not because vendors become more scrupulous. Washing persists while a label carries budget and lacks a boundary, so it fades when either changes. Two things are moving in that direction: buyers are getting better at asking for evidence rather than demonstrations, and the shared vocabulary around traces, evaluation, and governance is making capability easier to verify. Expect the term to lose its usefulness once buyers routinely ask what happens when the agent fails, since that question alone sorts most of the market.

Ask for the trace
Show me the trace of a run that failed.

It is the question that sorts the market, and it is a fair one to ask us. CAMS traces every model call and tool call per agent, gates promotion through Dev, QA and Production, and enforces spend ceilings and guardrails at runtime. Bring a workflow and we will run it.