all insights

Why AI agent projects fail: a field autopsy of the cancellation statistics

Gartner expects over 40% of agentic AI projects to be canceled by the end of 2027, and MIT found most pilots return nothing. Here's what actually kills these projects, and the pre-launch checks that predict survival.

Why AI agent projects fail: a field autopsy of the cancellation statistics — article illustration

In June 2025, Gartner predicted that more than 40% of agentic AI projects will be canceled by the end of 2027, citing “escalating costs, unclear business value or inadequate risk controls.” Two months later, MIT’s Project NANDA circulated a report claiming that 95% of enterprise GenAI pilots were producing zero measurable return. Both numbers went viral. Both are now quoted in board decks — usually stripped of every caveat their authors attached.

Here’s the uncomfortable part: the statistics are directionally right. We spend a lot of time inside Salesforce orgs where an agent pilot has stalled, and the causes are boringly consistent. It’s almost never the model. It’s the absence of a baseline metric, a consumption bill that grew faster than the value, an agent grounded on data nobody trusted, or a legal team that discovered the project three weeks before go-live.

This post does two things. First, it reads the cancellation statistics carefully — what Gartner, MIT, S&P Global, and RAND actually measured, and where the popular retellings go wrong. Then it walks through the seven failure modes we see in practice, and closes with the pre-launch checks that separate the agents that survive from the ones that become a line item in next year’s write-offs.

Four studies, four different autopsies

The failure statistics get lumped together as one big “AI doesn’t work” narrative, which is lazy. Each study measured something different, over a different window, with a different definition of failure. Put them side by side and the picture sharpens:

StudyPublishedHeadline numberWhat it actually measured
Gartner predictionJune 2025Over 40% of agentic AI projects canceled by end of 2027A forward-looking forecast, informed by a January 2025 poll of 3,412 webinar attendees
MIT Project NANDAAugust 202595% of GenAI pilots show no measurable P&L returnShort-window P&L impact of integrated pilots, from interviews, surveys, and 300+ public initiatives
S&P Global 451 ResearchMarch 202542% of companies abandoned most AI initiatives, up from 17% a year earlierSurvey of roughly 1,000 enterprises; the average org scrapped 46% of AI proofs of concept before production
RAND CorporationAugust 2024By some estimates, over 80% of AI projects fail — twice the rate of non-AI IT projectsInterviews with 65 experienced data scientists and engineers on root causes, mostly pre-agentic projects

Three details in the Gartner release deserve more attention than the headline. First, the same document predicts that by 2028, 15% of day-to-day work decisions will be made autonomously by agentic AI and a third of enterprise software will embed it — Gartner is forecasting a shakeout, not a collapse. Second, the poll behind it found only 19% of organizations had made significant agentic investments; 42% were being deliberately conservative. Third, Gartner estimates that of the thousands of vendors selling “agentic AI,” only about 130 offer real agentic capabilities. It calls the rest “agent washing” — chatbots, RPA, and assistants rebranded. Some fraction of the coming cancellations will be projects that were never actually agentic to begin with.

Gartner analyst Anushree Verma put the diagnosis plainly: most current projects are “early stage experiments or proof of concepts that are mostly driven by hype and are often misapplied.” She later expanded the argument in Harvard Business Review, and the thesis holds: the technology rewards organizations that approach it with discipline and strategic intent, and punishes the rest.

MIT’s “95% of pilots fail”: read the fine print before you quote it

The MIT number needs the most careful handling, because it’s the most quoted and the most misread. The report — The GenAI Divide: State of AI in Business 2025, from MIT’s Project NANDA — found that despite an estimated $30–40 billion in enterprise GenAI investment, 95% of organizations were getting zero measurable return from their integrated pilots. Only about 5% saw rapid, measurable value.

Before you build a strategy on that sentence, note the caveats:

  • “Zero return” meant P&L impact, over a short window. A pilot that improved response quality but hadn’t yet moved the income statement counted as a failure. ROI on process change routinely takes longer than a pilot cycle.
  • The methodology is described inconsistently across retellings. Coverage variously cites 150 interviews and a 350-person survey or 52 structured interviews and 153 surveyed leaders, alongside 300-plus public initiatives. When even the sample size is contested, treat the 95% as a warning light, not a law of nature.
  • The report’s own data undercuts the doom reading. Workers at the overwhelming majority of surveyed companies were quietly using personal AI tools daily — a “shadow AI economy” producing gains nobody was measuring. The failure was organizational absorption, not the technology.

The report’s genuinely useful findings got less airtime. Pilots built with specialized external partners succeeded roughly twice as often as internal builds. More than half of GenAI budgets went to sales and marketing tools, while the clearest returns showed up in unglamorous back-office automation. And the successful 5% shared a trait MIT called learning: their systems adapted to workflows and improved with feedback, rather than remaining generic tools bolted onto the side of a process.

So the honest synthesis across all four studies isn’t “agents don’t work.” It’s narrower and more actionable: most organizations deploy agents in a way that makes value unmeasurable, unaffordable, or ungovernable. That’s a fixable set of problems, which brings us to the taxonomy.

No baseline to beat, and consumption costs that outrun value

The first two failure modes are economic twins, and between them they cover most of Gartner’s “unclear business value” and “escalating costs.”

Failure mode 1: no baseline metric to beat. Ask a stalled project team what number the agent was supposed to move, and you’ll often get silence. Nobody recorded the pre-agent cost per case, deflection rate, or handle time — so six months in, nobody can prove the agent changed anything. RAND’s interviews found the leading root cause of AI project failure is stakeholders misunderstanding or miscommunicating the problem the AI is meant to solve. An agent with no baseline is unfalsifiable, and unfalsifiable projects die in budget reviews because “trust us, it’s helping” loses to any line item with a number attached. The fix costs almost nothing: measure the process for a month before the agent touches it, and write down the target it must beat. If you want a structured way to frame that business case, our Agentforce ROI calculator forces exactly those inputs.

Failure mode 2: consumption costs that scale faster than value. Agentic pricing is metered. On Salesforce, as of this writing, Flex Credits run $500 per 100,000 credits, with a standard agent action consuming 20 credits — about 10 cents per action — or a flat $2 per conversation on the alternative model. Every retrieval, record update, and tool call is an action. Now run the arithmetic on your own volumes: an agent that averages 25 actions per conversation costs roughly $2.50 per conversation on metered pricing. If the human-handled equivalent costs your business $4, the margin is real but thin — and it erodes every time the agent takes extra turns to disambiguate a vague request. Cost scales with activity; value scales with outcomes. Citi Ventures’ analysis of agentic pricing notes that consumption models bring exactly this problem: highly variable agent usage leads to unexpected expenses, which is why outcome-based pricing keeps coming up as the alternative. Until that matures, the discipline is yours: model cost per resolved outcome, not cost per month, and set alerts on actions-per-conversation drift. Projects that skip this discover their unit economics in the worst possible way — on an invoice.

Agents on broken data and briefs without borders

Failure mode 3: deploying agents over broken data. An agent is a very fast, very confident consumer of whatever your org contains. Duplicate accounts, stale knowledge articles, fields that mean different things in different business units — a human rep routes around that mess using tribal knowledge. An agent can’t. It grounds its answer in the wrong record and delivers it fluently. RAND lists lacking the necessary data as a top-five root cause of AI project failure, and in our experience it’s the one that surfaces latest, because demos run on clean sample data. In production the agent inherits a decade of deferred data hygiene on day one. If your knowledge base wouldn’t survive a new-hire relying on it exclusively, it won’t survive an agent either.

Failure mode 4: scope too wide. The pitch deck says “an agent for customer service.” The winning implementation says “an agent for order-status and return-eligibility questions, in English, for authenticated customers.” Verma’s sharpest line in the Gartner release was that many use cases positioned as agentic today don’t require agentic implementations at all — plenty of them are workflows wearing a costume. And breadth carries a mathematical penalty. Agent reliability compounds per step: a system that gets each step right 95% of the time completes a ten-step task correctly only about 60% of the time. Wide scope means more topics, more tools, more steps — and an evaluation surface too large to test honestly. Narrow agents ship, prove a number, and earn their expansion. Wide agents demo well and die in UAT.

No risk controls, no production: governance and the pilot-to-production gap

Failure mode 5: no risk controls, so legal kills it. “Inadequate risk controls” is the third leg of Gartner’s cancellation triad, and it has a specific failure signature: the project runs for months as a technical effort, then compliance reviews it late and asks questions nobody prepared for. What can it say? What data does it expose? Where’s the audit trail? Who’s accountable when it acts? If the answers are shrugs, a cautious counsel is right to block launch — an agent that acts on records is a different risk class from a chatbot that drafts text. The teams that clear this gate designed for it from the start: PII handling, topic guardrails, logged interactions, and defined escalation paths. The platform tooling exists — Salesforce’s Einstein Trust Layer provides zero-data-retention with external models, toxicity scanning, and a full audit trail of every prompt and response — but tooling only helps if governance is in the project plan before the build, with legal in the room at scoping, not at sign-off.

Failure mode 6: the pilot-to-production gap. Pilots flatter agents. The people who built the agent run the demo, on inputs they chose, with a human reviewing every output. Production is adversarial: typos, mixed languages, ambiguous asks, angry customers, prompt injection attempts, API timeouts mid-task. The tail of weird inputs that never appeared in the pilot is precisely what the agent will face at 2 a.m. on a Saturday. This is why S&P Global found the average organization scrapping 46% of AI proofs of concept before production — the PoC bar and the production bar are different heights. The countermeasure is unglamorous: batch-test at scale before launch, including hostile and malformed inputs. On Agentforce that means Testing Center and Agentforce DX test specs — hundreds of scripted utterances asserting expected topic, actions, and response, wired into CI — plus Salesforce’s own guidance to baseline agent accuracy against human performance in a sandbox before anything reaches a customer. If your test plan is “we chatted with it and it seemed good,” you don’t have a test plan.

The absorption problem: nobody owns the agent after launch

The seventh failure mode is the quietest, and it kills projects that survived everything else. The agent launches, works, and then slowly degrades — because no one owns it.

Agents aren’t projects; they’re operations. Products change, policies change, the knowledge base drifts, an integration gets versioned, and the agent’s accuracy sags a little each month. Meanwhile the transcripts pile up unread. Who reviews them weekly? Who owns the escalation queue? Who retires a topic when the underlying process changes? In most stalled deployments we look at, the answer to all three is “the project team,” which disbanded at go-live.

This is the organizational reading of MIT’s findings. The successful 5% of pilots weren’t running better models — they had systems that learned from feedback and line managers empowered to adopt them, while the failures bought tools and never rewired the process around them. An agent that no human is coaching is an agent that’s quietly getting worse. S&P’s research adds the corollary: organizations with higher project failure rates also report more user resistance, and resistance feeds on exactly this kind of visible neglect. Budget the operating model — named owner, transcript review cadence, retraining loop, monthly cost-per-outcome report — as part of the build, not as an afterthought. In practice it’s about a fifth of the effort and most of the longevity.

Measurable, bounded, grounded, governed

Strip away the vendor noise and the seven failure modes reduce to four properties you can check before a single credit is consumed. Measurable: there’s a baseline number, recorded before launch, that the agent must beat — and a cost-per-outcome model, not a cost-per-month hope. Bounded: the scope is narrow enough to test exhaustively, and someone has asked whether this use case needs an agent at all. Grounded: the data and knowledge the agent will stand on has been audited by someone who’ll sign their name to it. Governed: legal saw the design early, the audit trail exists, escalation paths are defined, and a named owner runs the agent after go-live.

Every canceled project we’ve examined failed at least one of these checks, and usually the team could have known on day one. That’s the real lesson of the statistics. Gartner’s 40% and MIT’s 95% aren’t verdicts on the technology — the same reports forecast agents making 15% of daily work decisions by 2028 and document a shadow workforce already using AI to real effect. They’re verdicts on deployment practice during a hype cycle, which means the failure rate is substantially within your control. Run the four checks honestly before you build, and you’re no longer betting on the industry average. You’re betting on your own preparation — much better odds.

Understanding the basics

Why do most AI agent projects fail?

Rarely because of the model. The recurring causes are economic and organizational: no baseline metric to prove value against, consumption costs that scale with activity rather than outcomes, agents grounded on poor-quality data, scope too broad to test, missing risk controls that trigger late legal blocks, untested adversarial conditions in production, and no owner for the agent after launch.

What did Gartner actually predict about agentic AI?

In June 2025, Gartner predicted over 40% of agentic AI projects would be canceled by the end of 2027 due to escalating costs, unclear business value, or inadequate risk controls. The same release estimated only about 130 of thousands of “agentic AI” vendors are genuinely agentic, yet also forecast that 15% of day-to-day work decisions will be made autonomously by agentic AI by 2028 — a shakeout, not a collapse.

Is the MIT statistic that 95% of AI pilots fail accurate?

Treat it as a warning, not a law. MIT Project NANDA’s 2025 report found 95% of integrated GenAI pilots showed no measurable P&L return, but the window was short, the methodology is described inconsistently across retellings, and the same report found widespread unmeasured productivity from employees’ personal AI use. Its stronger finding: pilots with external partners and narrow, learning-oriented scope succeeded roughly twice as often as internal builds.


Weighing an agent project against these failure modes — or trying to rescue one that’s already stalled? Talk to us — post-mortems and pre-mortems are both things we do.

Keep reading

All insights