Papers  /  MIT NANDA · the 95% number
Industry report~11 min readRead Aug 2026
Breakdown

The 95% number,
and the word carrying it

It is the most-quoted statistic in enterprise AI: roughly 95% of pilots produced no measurable profit-and-loss impact.17 It appears in board decks, in sales objections, and in every article about the AI bubble. It is also routinely misread in both directions — as proof the technology does not work, and as a problem someone else has. Read carefully, it is neither.

01

What the finding actually says

The claim, stated precisely

A study attributed to MIT’s NANDA initiative reported that of the enterprise generative-AI pilots it examined, roughly 95% produced no measurable impact on profit and loss. The figure has been widely reported in press coverage.17 reported via press coverage rather than a public paper — treat the exact figure as directional

Two words decide what this supports. Measurable: the finding is about attributable, demonstrated effect on a financial statement — not about whether users liked the tool, whether the model was accurate, or whether anyone felt more productive. And pilot: a bounded trial, not a production deployment, which selects for exactly the projects most likely to end before an accounting period closes.

→ The most common misreading

“95% of AI projects fail” is not what it says, and repeating it that way will lose you an argument with anyone who has read the coverage. A project can deliver real value and still show no measurable P&L impact — because nobody captured a baseline, because the effect was distributed across teams, or because it was never mapped to a line anyone reports. That distinction is the whole subject of this page.

02

Three separate failures hiding inside one number

They need different fixes, and conflating them is why the statistic gets quoted without changing anything.

FailureWhat actually happenedThe fix
Measurement failureThe system worked and nobody can prove it — no baseline was captured before the change, so the before-picture does not exist.Capture the baseline in week one, before you build anything. It is unrecoverable afterwards.
Attribution failureSomething improved, and three other things changed in the same quarter. The finance team will not credit yours.Agree one number, with one owner, in their reporting language, before the pilot starts.
Actual failureThe workflow was wrong, adoption never happened, or the system solved a problem nobody was paid to care about.Scoping. The problem was chosen badly, and no amount of engineering repairs that.

The first two are the reason a forward deployed engineer exists as a role. Both are solved by work that happens before and after the model — discovery that finds a workflow with a number attached, and instrumentation that makes the change attributable. Neither is model work, and neither gets done by shipping an API.

week 1the only time you can capture a baseline
one numberagreed, owned, in their language
~95%reported share with no measurable P&L impact17
03

Why this is the commercial case for the role

Read alongside a16z’s services-led-growth argument, the picture is coherent rather than contradictory. a16z’s thesis is that companies should deliberately accept lower gross margin on deployment work because the resulting customer outcome is defensible.8 The NANDA-style finding is what the market looks like when nobody does that work: capability is available, outcomes are not, and the gap between the two is a services gap rather than a model gap.

What the number supports

  • Enterprise AI value is bottlenecked on deployment and measurement, not on model capability.
  • Most organisations cannot currently prove the value of work they have already done.
  • A role that owns the arc from problem selection to demonstrated outcome has an obvious economic justification.
  • “It works” and “we can show it worked” are different projects, and the second is usually unstaffed.

What it does not support

  • That the technology does not work. The finding is about attribution, not capability.
  • That 95% of production deployments fail. It concerns pilots.
  • A precise figure. It is reported via press coverage rather than a public methodology, so quote it as approximate.
  • Fatalism. Every failure mode named above is preventable by work that costs days, not quarters.
→ Using it in a customer conversation

Bringing this number up yourself is a strong move, and only in one form: “most pilots never establish what would count as working, so let us agree that first — what number, measured how, and who signs off that it moved?” That reframes you from a vendor claiming an exception into the person solving the failure they have probably already lived through.

04

What it changes about how you work

1

Capture the baseline before you build

How many cases per week, how long each takes, what the error rate is today. In week one, while it is still available — after you change the workflow, the before-picture is gone permanently and every later conversation is anecdote against anecdote.

2

Agree one number, and who owns it

Not four metrics. One, in their reporting language, with a named person who will confirm it moved. The full method is in proving value & ROI.

3

Instrument adoption, not just accuracy

A system with excellent evals and no users produces exactly the outcome this statistic describes. Workflow DAU, override rate and deflection rate tell you whether anything actually changed.

4

Set the quality bar against humans

“Measurable” requires a comparison. Inter-expert agreement gives you a defensible one, and it also makes the target achievable — see evals for client work.

5

Say so when a deployment should not renew

Some pilots in that 95% were kept alive by reluctance to call them. An FDE who says it early, with evidence, is more valuable than one who defends a doomed account for two more quarters.

Practise the part this number is about

Measurement, not modelling

Frequently asked

Quick answers

What is the MIT NANDA 95% statistic?

A widely reported finding attributed to MIT’s NANDA initiative that roughly 95% of the enterprise generative-AI pilots it examined produced no measurable impact on profit and loss. It has been reported through press coverage rather than a broadly circulated public methodology, so the exact figure is best treated as directional rather than precise.

Does the 95% figure mean AI does not work?

No. The finding is about measurable, attributable financial impact, not about model capability. A project can deliver real value and still register nothing measurable — because no baseline was captured before the change, because the effect was distributed across teams, or because it was never mapped to a line anyone reports. It is a measurement and attribution finding first.

Why do most enterprise AI pilots show no measurable impact?

Three separate failures hide inside one number. Measurement failure: the system worked and no baseline exists to prove it. Attribution failure: something improved but three other things changed the same quarter, so finance will not credit yours. Actual failure: the workflow was wrong or nobody adopted it. The first two are preventable in week one; the third is a scoping problem.

How should a forward deployed engineer respond to this statistic?

Raise it yourself, in one specific form: most pilots never establish what would count as working, so agree that first — what number, measured how, and who signs off that it moved. Then capture the baseline in week one before anything changes, instrument adoption rather than only accuracy, and set the quality bar against measured human agreement rather than a round figure.

How does this relate to the argument for forward deployed engineers?

Directly. If capability is available but outcomes are not, the gap is a deployment and measurement gap rather than a model gap — which is the economic justification for a role that owns the arc from problem selection through to demonstrated outcome. It is also the mirror image of the services-led-growth thesis, which argues for deliberately spending margin on exactly that work.

Receipts

Sources

Every number on this page traces to one of these. Where a figure is self-reported or crowd-sourced rather than first-party, it is labelled inline.

Cited on this page

  1. TechCrunch — “Forward-deployed engineers are the AI industry’s latest talent obsession” (30 Jul 2026)techcrunch.com/2026/07/30/forward-deployed-engineers-are-the-ai-industrys-latest-talent-obsession/
  2. a16z — Services-Led Growth — the margin-for-moat thesis behind FDE hiring. a16z.com/services-led-growth/
  3. MIT NANDA “State of AI in Business” — 95% of pilots with no measurable P&L impact — widely reported; figure cited via press coverage rather than a public PDF. techcrunch.com/2026/07/30/forward-deployed-engineers-are-the-ai-industrys-latest-talent-obsession/
The 95% number, examined · part of the FDE track · Vibe Engines · 2026
Finished this one? 0 / 111 Paper Breakdowns done

Explore the topic

See this alongside everything else on the same subject — handbooks, system designs, challenges and tools, in one place.

More Paper Breakdowns