Automation · 10 min read

“95% of AI pilots fail”: unpacking MIT's 52 interviews and four other studies

“95% of AI pilots fail” was stitched together from a preliminary MIT NANDA report in which success meant whatever respondents called success. Experiments that measured the work show a different picture: gains for some workers and tasks, nothing or a loss for others.

An unlabelled bar chart: the tallest bar propped up by a thin stick, and a magnifying glass with a highlighted rim held to its base

Where the 95% came from

In August 2025 the business press ran with the line “95% of AI pilots fail”. The source was an MIT NANDA report, “The GenAI Divide: State of AI in Business 2025”, dated July 2025 on its title page. The file is marked v0.1, and the second page reads Preliminary Findings. The listed reviewer is one of the co-authors; there was no external peer review. MIT no longer hosts the report, so we read a copy on a third-party site.

The executive summary says: “Despite $30–40 billion in enterprise investment into GenAI, this report uncovers a surprising result in that 95% of organizations are getting zero return” (MIT NANDA, 2025, v0.1 copy). In the next sentence the authors add that only 5% of integrated AI pilots bring in millions, and the vast majority show no measurable effect on P&L. Those two sentences became “95% of pilots fail”. But a pilot that does not bring in millions has not necessarily failed.

Now the method. The authors reviewed more than 300 publicly disclosed AI initiatives, held structured interviews with representatives of 52 organisations, and collected 153 survey responses from senior leaders at four industry conferences. The research period was January to June 2025.

A task-specific tool counted as a success if users or executives said it had produced a noticeable, lasting effect on productivity or P&L. Said, not shown in the accounts. The authors flag this themselves: the figures are directionally accurate, rest on interviews rather than official company reporting, and definitions of success may differ between organisations.

The report has a second number, about enterprise AI systems, custom-built or bought from a vendor. 60% of organisations evaluated them, 20% got as far as a pilot, 5% reached production. That is a different claim, and it covers only those systems: ChatGPT and Copilot are not in it. Read the funnel as drawn, and one in four organisations that ran a pilot made it to production, not one in twenty.

On ChatGPT and Copilot, the same report says they are widely adopted: over 80% of organisations have explored or piloted them, and nearly 40% report deployment. 40% of companies bought an official LLM subscription, while workers at over 90% of companies regularly use personal AI tools for work. In the authors’ view, these tools mainly raise individual productivity rather than P&L. The McKinsey survey below shows the same gap.

52 + 153
interviews with organisations and conference surveys — the basis of the MIT NANDA report
−19 pp
correct solutions among consultants using GPT-4 on a task beyond the “frontier” of AI capability
37%
of McKinsey respondents in 2026 see at least some EBIT impact from AI — about the same as in 2025

Call centre: the novices gained

Erik Brynjolfsson, Danielle Li and Lindsey Raymond studied the customer support operation of a US software company. The channel was text chat. Agents got a GPT-based assistant that suggested replies; a human still sent the answer.

The metric was issues resolved per hour. In the NBER working paper of April 2023, covering 5,179 agents, the average gain was 14%: 34% for novice and low-skilled workers, minimal for experienced and highly skilled ones. The version published in the Quarterly Journal of Economics in 2025 uses a refined sample of 5,172. There the average effect is 15%, and around 30% for less skilled and less experienced workers. The most skilled got slightly faster and slightly worse on quality. On top of that, customers wrote to the chat in a friendlier tone, and agents quit less often.

The authors’ explanation is that the tool passes the practices of the best agents on to novices. Judging by the numbers, the experienced agents had little left to pick up.

For a pilot, the lesson is simple. An average of 15% is made up of a clear gain for some people and almost none for others. Count only the average, and you cannot see who needs the tool and who it gets in the way of.

BCG: the frontier beyond which AI hurts

In 2023 researchers from HBS, Wharton and MIT Sloan ran an experiment with BCG on 758 consultants, about 7% of the firm’s individual-contributor consultants. The design was pre-registered. The tool was GPT-4.

On 18 tasks within the model’s capabilities, consultants with AI completed 12.2% more tasks, worked 25.1% faster, and produced work of more than 40% higher quality. Those who had scored below average before the experiment improved by 43%, those above average by 17%. This is the part that usually gets retold.

One task was deliberately chosen to sit outside what the model could do. On it, consultants using AI “were 19 percentage points less likely to produce correct solutions compared to those without AI” (HBS Working Paper 24-013, 2023). The paper’s title calls this frontier jagged.

The study has two limits. The tasks resemble consulting work but were written for the experiment. And the model is the 2023 GPT-4. According to one of the authors, the paper came out in Organization Science in 2026, but we have not read the final text and cite the working paper.

METR: slower in 2025, unclear in 2026

In July 2025 METR published a randomised controlled trial: 16 experienced developers from large open-source projects, 246 real issues, Cursor Pro with Claude 3.5 and 3.7 Sonnet. Beforehand the developers expected a 24% speed-up. Afterwards they believed they had been 20% faster. As measured, tasks with AI took 19% longer.

METR said straight away that this was a snapshot of early 2025 in one setting, not proof that AI is useless to most developers.

In February 2026 came an update covering late-2025 tools. For the original group, time per task was −18%, meaning a speed-up, but with a confidence interval from −38% to +9%. For newly recruited developers it was −4%, with an interval from −15% to +9%. Both intervals include zero: there is no statistically significant effect in either direction.

The reason for changing the design is more telling than the figures. Participants are paid $50 an hour, and still a growing share of developers do not want to do half their work without AI. The people keenest on AI drop out of the sample, so by METR’s estimate the real speed-up may be considerably higher than measured. “Due to the severity of these selection effects, we are working on changes to the design of our study,” METR writes (February 2026). The page for the 2025 study now carries a note that those results no longer reflect the current effect.

Anyone quoting “19% slower” in 2026 without this update is repeating an outdated measurement.

The gap Solow described

On 25 August 2026 The Register reported on a new McKinsey survey, “The state of AI in 2026: On the road to ROI”. It covers 1,719 respondents. Eight in ten say AI has improved their own productivity. 37% attribute at least some EBIT impact to AI, about the same as a year earlier. Around 6% attribute at least 5% of EBIT to AI and call the impact significant.

Both numbers are what respondents said, not figures from company accounts. But the questions differ: the 80% is about people’s own work, the 37% is about company profit. The first does not turn into the second on its own.

The gap is old. In 1987 the economist Robert Solow wrote in a review for the New York Times Book Review: “You can see the computer age everywhere but in the productivity statistics” (as cited by Yoram Bauman). Conversations about the productivity paradox usually start from that line. Thirty-nine years later it needs barely an edit: eight in ten feel more productive, and 37% see it in EBIT.

What to measure in your own pilot

The three experiments above differ from the MIT report in one respect: they measured the result, while MIT asked about it. For a pilot on one process, that turns into six rules.

  1. Write the metric down before the start. Researchers call this pre-registration, and the BCG experiment was run that way: the hypothesis and the method of measurement are fixed before any data exists. For a pilot, one sentence is enough: what we count, in what units, and what number means success.
  2. Measure “before”. Work as usual for a week and count the same thing. If the volume allows, a control group is better: part of the enquiries keep going the old way.
  3. Count novices and experienced staff separately. The average hides who actually gained.
  4. Measure quality alongside speed. Beyond the model’s frontier, BCG consultants got more answers wrong, and the strongest call-centre agents lost a little quality.
  5. Time it; don’t ask. METR’s developers were sure they had sped up, and the measurement showed a slowdown.
  6. Convert it to money. 80% of McKinsey’s respondents report higher personal productivity; 37% see an effect on EBIT. A pilot has to answer the second question.

How to choose the process for such a pilot is covered in our piece on what to automate first.

Our pilots run for two weeks, and we write down the success metric before work starts. It is the same pre-registration, scaled down to one process. We have no results from our own pilots yet, which is why there is not a single number of ours in this article.

What we don’t know

The MIT NANDA report is version v0.1, marked Preliminary Findings, with no external peer review. MIT no longer hosts it; we read a copy on a third-party site. The McKinsey page would not open for us: the figures come from The Register’s coverage, and we have not checked the survey dates or the exact wording. We have not seen the final BCG paper in Organization Science. The Solow quote is given as cited by Bauman; we did not open the New York Times archive.

The call centre is one company and one type of task. The BCG experiment used tasks written for the study and a 2023 model. None of these studies was run on small businesses in Uzbekistan or the UAE, and their percentages do not carry over to your process directly.

How we run such pilots — one process at the start and a baseline metric in the contract — is described on AI implementation for business.

If you want someone from outside to break down a process for a pilot like this, describe it in the questionnaire on our home page. The breakdown of one process arrives within 48 hours, free and without a call.

Read next

Show your process — I will send back an automation map in 48 hours

Six questions about how things are set up at your place today. The output is a diagram: what can be taken off people, in what order and what it costs.