Why 95% of AI pilots fail — and what the 5% do differently
MIT's GenAI Divide study became 2025's most quoted AI statistic. Here is what it actually measured, what the media got wrong, and the pattern behind the projects that worked.

In August 2025, an MIT-affiliated research group published a report with a number that moved markets: 95% of enterprise generative AI pilots delivered no measurable P&L impact. It got quoted in boardrooms, bear cases, and roughly ten thousand LinkedIn posts.
Most of the people quoting it never read it. That is a shame, because the useful part of the study was never the headline — it was the anatomy of the 5% that worked.
What the number actually measured
The study — MIT Project NANDA's GenAI Divide: State of AI in Business 2025 — analyzed roughly 300 enterprise AI projects. "Failure" did not mean the software broke. It meant the project produced no marked, sustained profit-and-loss impact. That is a very high bar, and an important distinction:
- A pilot that people liked but nobody measured? Counted as failure.
- A tool that saved time that never showed up in a budget line? Failure.
- A technically flawless deployment of the wrong workflow? Failure.
The divide was not model quality. It was whether the system was wired into a real workflow with someone accountable for the outcome.
The three findings worth stealing
~2x
success rate for purchased or partner-built systems vs. internal builds
Specialized vendors ship the integration lessons they learned elsewhere
90%
of employees already use personal AI tools at work
The shadow AI economy captures value official pilots miss
5%
of integrated pilots extracted significant value
Narrow scope, deep workflow integration, measured baselines
First: buy-and-adapt beat build-from-scratch, roughly two to one. Internal teams kept rebuilding infrastructure that already existed instead of spending that effort on workflow fit — which is the part that actually determines success.
Second: shadow AI is real adoption data. While official pilots stalled, employees at nine out of ten companies were quietly using personal ChatGPT accounts to do their jobs. That is not a compliance anecdote; it is a map of where the demand actually is.
Third: the winners picked back-office drudgery, not moonshots. Document processing, reporting, triage, chase-ups. Unsexy workflows with clear before/after numbers.
Why pilots die: the learning gap
The study's core diagnosis is what it calls the learning gap. Generic tools do not remember your context, do not adapt to your process, and break the first time reality deviates from the demo. A chat window is a capability; it is not a system.
The 5% closed the gap the same way:
- One workflow, fully wired. Trigger, inputs, actions, escalation — integrated with the system of record, not sitting beside it in a tab.
- A baseline before launch. Hours per week, cost per ticket, days to invoice. If there is no "before," there will never be a measurable "after."
- A named owner. Someone whose job includes tightening the system every week for the first month.
- Tolerance for friction. The successful teams treated early failure cases as tuning data, not as proof the project was doomed.
The honest caveat
Read carefully, the study is less pessimistic than the headline. Among companies that actually piloted custom tools, roughly a quarter reached successful deployment — a reasonable strike rate for any new technology. The 95% figure blends in companies that never seriously deployed anything. The lesson is not "AI does not work." It is that casual AI does not work.
A checklist for being in the 5%
- The workflow already exists and already hurts (daily or weekly, measurable)
- Success is a number on a report someone already reads
- The system writes into your tools, not into a separate dashboard
- A human escalation path exists and someone owns the queue
- You compared buying, adapting, and building before choosing
- There is a 30-day review cadence with authority to kill it
That checklist is deliberately dull. The GenAI divide is not between companies with better models — everyone rents the same models. It is between companies that run AI like a project and companies that run it like a press release.
If you want the tactical version of this — how to cut a workflow down to a shippable first slice — we wrote a scoping method for AI automation projects that pairs well with this one.


