Skip to main content

MIT · Project NANDA · State of AI in Business 2025

Why AI Projects Fail

Ninety-five percent of enterprise GenAI pilots return nothing to the bottom line. The cause isn’t model quality, regulation, or talent — it’s that ROI and operations were never engineered into the pilot.

95%return zero P&L impact
5%reach production
95%

of organizations see no P&L return on current pilots

5%

of integrated pilots reach sustained production

67%

deployment success when an external partner is involved

01 — The number everyone quoted

The headline was right. The reading was wrong.

Fortune, Forbes, and AOL all led with the same figure — “95% of GenAI pilots fail” — then stripped away every condition attached to it. The number is real. The conclusion most people drew from it isn’t.

What the coverage said

“AI doesn’t work.”

The failure framing dominated: zero return, broad disappointment, a technology that overpromised. Skeptics read confirmation; optimists dismissed it as early-stage noise.

What almost no headline carried were the positive correlates sitting in the same dataset — the conditions under which pilots actually succeed.

What the report found

The divide is about approach.

The gap between the 5% that scale and the 95% that stall isn’t driven by model quality or regulation. It’s driven by how the pilot was designed: whether it was embedded in real workflows, whether it learned over time, and whether anyone modelled the ROI before building.

Tools that don’t retain feedback, adapt to context, or improve get quietly abandoned — no matter how good the underlying model is.

Trust isn’t low because the models are bad. It’s low because ROI and operations were never engineered in.

02 — The playbook

Five ways to cross the divide.

Crossing isn’t a model decision — it’s an operating discipline. The organizations on the right side of the divide do the same five things, in roughly this order.

01

Design the pilot with the end in mind

Define success before you design anything. If you can’t write the thesis as a CFO-style equation — cost taken out, or revenue added — it isn’t a pilot yet, it’s an experiment looking for a justification.

The test: “Cut Tier-1 handling time 50% via AI triage.” “Automate invoicing to remove $5M of outsourcing.” Concrete, owned, measurable — or don’t start.
02

Embed it — don’t isolate it

AI that sits beside the workflow rarely scales; AI that lives inside the system of record does. Start in one contained process, pair generative output with deterministic rules where reliability matters, then expand.

67% vs 33% deployment success for embedded, partner-built systems versus internal builds. Mid-market teams reached production in ~90 days vs 9+ months for large enterprises.
03

Build the feedback loop into the product

Capture corrections, overrides, and edits as learning signals — in the interface, not in a monthly survey. Preserve memory across sessions. Learning velocity, not day-one accuracy, decides long-term value.

66% of executives now require systems that learn from feedback. 63% demand contextual memory in vendor selection. Pilots that don’t learn after launch almost never scale.
04

Model the ROI — prove value or pivot fast

Estimate costs and benefits before or during the pilot. Measure time saved, error reduction, throughput, and revenue — not demo counts. Killing a project that won’t deliver is a success: it frees resources for higher-yield work.

Back-office wins: $2–10M in annual BPO savings, 30% lower agency spend, $1M saved on outsourced risk. Gartner expects ~40% of agentic projects to be cancelled as ROI pressure rises.
05

Own it — accountability and change management

Someone must be accountable for outcomes in production. Decentralize who runs the deployment, but keep accountability tight, and let the people already using AI lead the rollout.

Find your prosumers. Employees already using ChatGPT or Claude make the best champions. Externally built tools saw nearly 2× the usage of internal builds.
Bonus — what separates leaders

It isn’t budget or model choice. It’s foundational data readiness.

the revenue gains of peers

the cost savings of peers

03 — The Monday-morning takeaway

Five rules for the CEO.

Rule 01

If you can’t express the thesis as a CFO-style equation, it isn’t a pilot yet.

Rule 02

For high-stakes work, a deterministic spine with AI at the edges beats pure generative autonomy.

Rule 03

If it isn’t embedded in the system of record, adoption will be cosmetic.

Rule 04

Build feedback capture into the UI, not into a monthly survey.

Rule 05

Start with your data health first, then bolt GenAI on from there.

And then

Find your prosumers

The thesis

This isn’t about being an AI optimist or pessimist. It’s about being operationally serious.

04 — The conversation

Hear the full breakdown.

We went deep on the playbook — pilot design, integration, feedback loops, ROI, and ownership. Press play and the episode picks up exactly where the practical playbook begins.

Jumps to 1:01:00 — the playbook

Prefer to start from the top? Once it’s playing, drag the scrubber back to 0:00 — only the entry point is set to the playbook.

05 — The research

Read the source.

Every figure on this page comes from MIT Project NANDA’s 2025 study of enterprise AI implementation. Download the full report and check the numbers yourself.

Download the PDFThe GenAI Divide · 26 pp · ~0.9 MB

Figures sourced from MIT Project NANDA, “The GenAI Divide: State of AI in Business 2025.” Commentary and playbook framing are iSolutionsAI’s own. Gartner projection cited where noted.

Ready to be in the 5%?

We engineer ROI and operations into AI from day one — not as an afterthought. Let's build a pilot that ships and sticks.