Note

Most AI pilots fail. The reason is boring.

MIT studied more than 300 enterprise AI deployments and found 95% produced no measurable return. The failures had almost nothing to do with the models.

In mid-2025, researchers at MIT published a study of more than 300 enterprise AI deployments. The headline finding: 95% of generative AI pilots produced no measurable return. Not modest returns — none that the companies could find.

A year on, the number is still traveling, mostly as ammunition. Skeptics quote it as proof the technology is hype; vendors quote it as proof that everyone else is doing it wrong. The useful material is in the middle of the report, in the anatomy of why the pilots failed.

It was not the models. The report’s name for the core problem is the learning gap: most AI tools, as deployed, “do not retain feedback, adapt to context, or improve over time.” One executive put it plainly — the tool was excellent for brainstorming and first drafts, but it did not remember client preferences or learn from previous edits. It performed beautifully in the demo, then performed identically forever, no matter what anyone taught it.

That is the difference between a demonstration and a system. A demonstration has to be impressive once. A system has to be right on an ordinary Tuesday, and when it is wrong, something specific has to happen next. Almost none of the failed pilots had an answer to the Tuesday question — no measurement of whether the output was correct, no defined step for when it was not, no way to tell whether last month’s changes made things better or worse.

Two quieter findings deserve more attention than the headline number.

First, the money went to the wrong end of the business. More than half of the AI budgets in the study went to sales and marketing, because that is where the exciting demonstrations live. The measurable returns showed up in back-office work — document handling, intake, processing. The reading, sorting, and re-typing that nobody puts on stage turned out to be where the money was.

Second, who built it mattered enormously. Systems built with outside specialists reached successful deployment about 67% of the time. Internal efforts succeeded about 33% of the time. Building production AI turned out to be a specialty rather than a side project, which is roughly what happens whenever specialized work is attempted as a sideline.

None of this says the technology does not work. The same report found the successful 5% extracting millions in value, and found that roughly 90% of employees were already using AI tools personally whether or not their employer had bought anything. The technology is fine. What fails is the pattern: buy something general, point it at nothing in particular, measure nothing, and call it a strategy.

The pattern in the successful few is not ambition. It is specificity — one process, understood in detail; a person kept in the loop wherever a wrong answer is expensive; measurement in place from the first week. A pilot that cannot answer the question “how will we know it is working” has already answered it.

  • ai
  • operations

Get in touch

Tell us what you need built.

We take on a few projects at a time. If it is not a fit, we will say so quickly.