In mid-2025, MIT's NANDA initiative published a study that landed harder than most enterprise technology research ever does, largely because of a single number that contradicted essentially every vendor deck circulating at the time. The number was 95%, and what it described was the share of enterprise generative AI pilots producing no measurable impact on the profit and loss statement at all. For a Canadian business owner being told weekly that AI adoption is now existential, the study is worth understanding precisely, including the parts that complicate its own headline.

Key Takeaway

MIT's NANDA initiative published The GenAI Divide: State of AI in Business 2025, finding that despite an estimated US$30 to 40 billion in enterprise generative AI investment, roughly 95% of organizations studied were getting zero measurable return, with only about 5% of pilots delivering significant value. The study's own diagnosis was not that the models were inadequate; it identified a "learning gap," tools that cannot retain feedback, adapt to context, or improve over time, combined with poor integration into actual workflows. Two findings matter most for finance specifically: over half of AI budgets went to sales and marketing despite the study finding better returns in back-office operations and finance, and externally-sourced tools succeeded roughly twice as often as internally-built ones. The number deserves scrutiny, its sample and methodology are modest and the "ROI" framing has been credibly criticized, but the underlying diagnosis is well-supported and directly actionable.

The Study, And What It Actually Measured

The research, published by MIT's Media Lab through its NANDA initiative under the title The GenAI Divide: State of AI in Business 2025, drew on approximately 150 interviews with business leaders, a survey of roughly 350 employees, and analysis of 300 publicly documented AI deployments[1]. Its stated objective was to understand why a small number of organizations achieve significant results from generative AI while the majority remain stuck with unfulfilled promises[2].

The specific measure matters and is frequently misreported. The study assessed whether pilots produced measurable impact on revenue or the P&L, not whether they were technically functional, whether employees liked them, or whether they produced any benefit at all. A pilot that saved staff time without that time translating into a measurable financial outcome would register as failure under this definition. This is a demanding standard, and it is worth holding in mind when the 95% figure gets deployed rhetorically as evidence that AI "doesn't work," which is not what the study measured or concluded.

The Headline Number, In Context

The specific findings are worth citing precisely. Despite an estimated US$30 to 40 billion in enterprise investment into generative AI, the report found approximately 95% of organizations were getting zero return, with only about 5% of AI pilots achieving rapid revenue acceleration[3]. The deployment funnel was similarly stark: of firms evaluating enterprise-grade systems, roughly 60% evaluated them, about 20% reached pilot stage, and only around 5% went live[4].

The report also found sectoral concentration in where transformation was actually occurring: only two of nine major sectors examined, technology and media, showed material business transformation from generative AI use[4]. And it identified what the authors termed an enterprise paradox, that large firms led in pilot volume while lagging in successful deployment[4], a finding with some genuinely encouraging implications for smaller businesses discussed below.

The Learning Gap

The study's central explanatory concept, and the part most useful to a business planning its own adoption, is what the authors call the learning gap. Their diagnosis was explicitly not that model quality was the constraint. As the report's authors put it, pilots stall because most tools cannot retain feedback, adapt to context, or improve over time[5]. Lead author Aditya Challapally described the mechanism directly in subsequent interviews: generic tools like ChatGPT excel for individuals precisely because of their flexibility, but stall in enterprise use because they do not learn from or adapt to specific organizational workflows[3].

The report illustrated this with a Fortune 500 insurer whose sanctioned generative AI pilot looked polished in boardroom demonstrations but collapsed in field use because it could not retain context[5]. This is a specific, recognizable failure pattern: a tool that performs impressively on a curated demonstration task and then fails to survive contact with the messy, context-dependent, institution-specific reality of actual work. A manufacturing COO quoted in the research captured the resulting disconnect bluntly, describing a gap between what industry commentary claimed had changed and what had actually shifted in their operations[4].

The Budget Misallocation Finding

This is the finding with the most direct relevance to a finance function, and it deserves emphasis because it inverts where most organizations were actually spending. The study found AI budgets overwhelmingly favoured sales and marketing applications, with over 50% of 2025 AI budgets directed there, despite the research identifying better returns in back-office operations and finance[4][6]. Independent analysis of the same findings characterized this as a mismatch of use case and value: the highest returns lay in back-office automation, document processing, compliance, internal workflows, while the money flowed to higher-visibility, lower-ROI front-office pilots[7].

The explanation for this misallocation is not mysterious, and it connects directly to behavioural patterns this publication has examined elsewhere. Sales and marketing AI pilots are visible, demonstrable, and produce compelling narratives for a board. Back-office finance automation is unglamorous, harder to showcase, and produces savings that accrue quietly across many small processes rather than as a single attributable win. The vividness asymmetry that drives discount creep in pricing decisions operates here too, on capital allocation for technology.

Build Versus Buy: The 67% Finding

A second finding with immediate practical consequence concerns sourcing. The research found that tools built by external vendors succeeded roughly twice as often as internally-built systems[4]. A related analysis of the same data reported that pilots blending internal AI specialists with external expertise achieved a 67% success rate versus only 22% for IT-only internal builds[6].

This finding cuts against a common instinct, particularly at organizations with capable technical teams, that building internally provides better fit and avoids vendor dependency. The study's data suggests the opposite in practice, and the plausible mechanism connects back to the learning gap: an external vendor building the same category of tool across many customers accumulates domain-specific learning and workflow integration experience that a single internal team building one instance cannot match. For a Canadian SMB without an internal AI team at all, this finding is straightforwardly good news, since it suggests the resource disadvantage is smaller than it appears.

The Methodological Caveat Worth Stating

Intellectual honesty requires noting the study's own limitations rather than treating a striking number as settled fact. The sample is modest: roughly 150 executive interviews, 350 employee surveys, and 300 documented deployments[1], and reporting of these figures varies slightly across secondary coverage, with some sources citing 52 executive interviews and 153 leader surveys[8], a discrepancy in the secondary literature worth noting since it suggests coverage has not been uniformly careful with the underlying numbers.

This is also a single study from a single research initiative, published outside the peer-reviewed academic literature, that received unusually heavy media amplification. Its central findings are plausible and consistent with other research, Gartner analysis cited alongside it found over 40% of agentic AI projects stalling[9], but a business making a substantial strategic decision should treat the 95% figure as a well-publicized directional indicator from one credible source rather than an established empirical constant.

The "Wrong Metric" Critique

A substantive intellectual challenge to the study deserves airing, because it is credible and changes how the finding should be applied. A response paper from a Berkeley executive education faculty director argued the headline statistic reveals a more fundamental question than it answers: whether traditional ROI is the right metric for AI implementations at all, and whether the expectation that a new general-purpose technology should accrue immediate financial benefits reflects the same flawed thinking that has accompanied every major technological transformation[10].

The critique has real force. Measured on immediate P&L impact, early enterprise adoption of spreadsheets, email, or the web would likely have registered similar failure rates, and the eventual value of those technologies was not visible in the first eighteen months of adoption. This argument should not be used to excuse genuinely wasted spending, a pilot that has produced nothing and has no articulable path to producing anything is a failed pilot regardless of framing, but it does argue against treating a two-year P&L test as the sole criterion for whether AI adoption is working, particularly for capability-building investments whose payoff is deliberately longer-dated.

What The 5% Actually Do Differently

The study and subsequent analyses converge on a consistent profile for the minority of pilots that succeeded, and it is refreshingly unexotic. Successful initiatives were tightly scoped rather than broad, domain-specific rather than general-purpose, and sourced through smart external partnerships rather than built in isolation[6]. Lead author Challapally described the successful startups in the dataset in similar terms: they pick one pain point, execute well, and partner smartly with the companies who use their tools[3].

One analysis added a counterintuitive framing worth considering: that the 5% succeed by designing for friction, human, organizational, and technical, rather than attempting to eliminate it, on the argument that the resistance a deployment encounters is itself the information that drives adaptation[5]. Whether or not one accepts that specific framing, the underlying observation, that pilots which never encounter or absorb real operational resistance tend to be the ones that look good in demos and fail in the field, is consistent with the Fortune 500 insurer example the study itself documented.

The 18-Month Window Claim, Examined

One assertion circulating alongside the MIT findings deserves specific scrutiny because it is frequently used to manufacture urgency in vendor conversations. Analysis published alongside coverage of the study argued that early adopters gain competitive advantage through accumulated training data, shrinking the window for laggards, and that organizations have roughly 18 months to pivot to learning-capable systems before early adopters lock in advantages defining market position for the next decade[11].

This claim should be read carefully and largely discounted for most Canadian small and mid-sized businesses, for two reasons. First, it is an interpretive projection by a commercial analytics vendor rather than a finding of the MIT research itself, and the specific 18-month figure has no empirical derivation offered. Second, and more substantively, the data-accumulation advantage it describes applies most strongly to businesses whose competitive position depends on proprietary data scale, which describes very few professional services firms, distributors, manufacturers, or trades businesses. A plumbing contractor does not lose the decade because it adopted invoice automation in 2028 rather than 2026. Treating a vendor's urgency framing as an established research finding is exactly the kind of pressure that produces the hastily-scoped, poorly-defined pilots the MIT study found failing at 95%.

A Worked Case: The Pilot That Should Have Been Killed

A mid-sized Canadian professional services firm ran a generative AI pilot for client-facing proposal generation, a high-visibility use case championed by its business development lead. Eight months and roughly $60,000 in software and internal time later, the tool was producing proposals that required so much editing that the firm's proposal writers found it faster to start from their own templates. No proposal win rate improvement was measurable. The pilot continued anyway, on the reasoning that abandoning it would waste what had been invested.

Two things were true simultaneously, and separating them is the actual lesson. First, this was a textbook escalation of commitment, continuing to fund a failing initiative because of unrecoverable prior spending, a pattern this publication has examined at length. Second, and more specific to the MIT findings, the pilot exhibited the precise profile the study associates with failure: front-office rather than back-office, broadly scoped rather than tightly defined, and a general-purpose tool applied to a highly context-dependent task it could not retain organizational knowledge for.

When the firm eventually redirected the same budget to a narrowly scoped back-office application, automated extraction and coding of supplier invoices, it produced measurable savings within a quarter. Nothing about the firm's AI capability had changed. What changed was the use case selection, from the vivid, visible, front-office application to the unglamorous, back-office one the MIT research had identified as where the returns actually were.

Applying This At Canadian SMB Scale

Several of the study's findings translate unusually favourably to smaller businesses, which is worth stating because most coverage frames the research as a cautionary tale for enterprises without noting the asymmetry. The enterprise paradox finding, that large firms lead in pilot volume but lag in successful deployment[4], suggests organizational complexity is itself an obstacle, and a smaller business has less of it. The build-versus-buy finding means the absence of an internal AI team is closer to an advantage than a handicap. And the back-office concentration of returns means the highest-value applications, document processing, compliance workflows, invoice and expense automation, are precisely the ones available to a small business through commercial tooling rather than requiring bespoke development.

The practical translation is a specific sequencing recommendation: start in the back office, not the front; scope to a single, well-defined, repetitive process rather than a broad capability; buy rather than build; and define, in advance and in writing, what measurable outcome would constitute success and by when, which is the pre-commitment discipline that makes the difference between a pilot that gets honestly evaluated and one that quietly persists for eight months past its usefulness.

An Honest Scorecard Before You Start

Drawing directly from the study's failure profile, a short set of questions worth answering in writing before funding any finance AI pilot. Is this back-office or front-office? The research locates returns in the former and spending in the latter. Is the scope one specific process, or a general capability? Tight scoping characterizes the successful minority. Are we buying or building? External sourcing succeeded at roughly twice the rate. What specific, measurable outcome would make this a success, and by what date? Defined in advance, this is the kill criterion that prevents the eight-month drift in the worked case above. Does the tool retain context and improve, or does it restart cold each time? This is the learning gap, stated as a procurement question.

The Findings At A Glance

For quick reference: estimated enterprise generative AI investment, US$30-40 billion. Share of organizations getting zero measurable return, approximately 95%. Share of pilots delivering significant value, approximately 5%. Deployment funnel: roughly 60% of firms evaluated enterprise systems, about 20% reached pilot, around 5% went live. Sectors showing material transformation: two of nine (technology and media). Share of 2025 AI budgets directed to sales and marketing: over 50%, despite returns concentrating in back-office operations. Success rate for pilots blending internal and external expertise: 67%, versus 22% for IT-only internal builds.

The Limits Of This Analysis

Several caveats matter. This article rests substantially on a single research initiative's findings, and while those findings are widely cited and consistent with adjacent research, they are not peer-reviewed and the sample sizes reported vary across secondary coverage, as noted above. The study measured a demanding criterion, measurable P&L impact, and the credible critique that this is the wrong metric for early general-purpose technology adoption is presented here rather than dismissed. The AI tooling landscape has also continued to develop since the study's mid-2025 publication, and specific findings about tool capability limitations, particularly around context retention, may be less true of current tooling than of the tools evaluated during the research period. What is most durable, and what this article recommends acting on, is the diagnostic pattern rather than the specific percentage: scope tightly, favour back-office applications, buy rather than build, and define success criteria before funding rather than after.

Frequently Asked Questions

What exactly did the MIT study find?
MIT's NANDA initiative published The GenAI Divide: State of AI in Business 2025, finding that despite an estimated US$30-40 billion in enterprise generative AI investment, roughly 95% of organizations were getting zero measurable return on the P&L, with only about 5% of pilots delivering significant value.
Does this mean AI doesn't work for finance?
No, and the study's own findings argue the opposite for finance specifically. The research located the highest returns in back-office automation, document processing, compliance, and internal workflows, while finding that over half of AI budgets went to lower-return sales and marketing applications instead. The problem the study identified was allocation and integration, not the technology's usefulness in finance.
Should we build our own AI tools or buy them?
The study found externally-sourced tools succeeded roughly twice as often as internal builds, with one analysis of the data reporting 67% success for pilots blending internal and external expertise versus 22% for IT-only internal builds. For a business without an internal AI team, this is genuinely encouraging.
Is the 95% figure reliable?
It should be treated as a credible directional finding rather than an established constant. It comes from a single, non-peer-reviewed study with a modest sample, and secondary coverage reports the sample sizes inconsistently. Its central diagnosis is consistent with adjacent research, including Gartner findings on stalled agentic AI projects.
What is the "learning gap"?
The study's term for its central diagnosis: pilots fail not because model quality is inadequate but because most tools cannot retain feedback, adapt to organizational context, or improve over time, so they perform well in curated demonstrations and fail when they meet the context-dependent reality of actual workflows.
What should a small business actually do first?
Start in the back office rather than the front, scope to one specific repetitive process rather than a broad capability, buy rather than build, and define in writing what measurable outcome would constitute success and by when, before funding the pilot rather than after it has drifted.
IB

About The Insight Bureau Research Desk

The Insight Bureau is GSH Financial's research publication, written for Canadian business owners and the students who will eventually advise them. This article examines a single influential study alongside credible published critiques of it; see References below.

References

  1. MIT Media Lab NANDA Initiative. (2025). The GenAI Divide: State of AI in Business 2025. Sample as reported: approximately 150 executive interviews, 350 employee surveys, and 300 documented public AI deployments.
  2. Santolo, F. (2026, March 7). 95% of Corporate Generative AI Projects Fail, MIT Study Finds. Medium. medium.com/@Fransantolo/95-of-corporate-generative-ai-projects-fail
  3. Fortune / Yahoo Finance. (2025). MIT Report: 95% of Generative AI Pilots at Companies Are Failing, including interview with lead author Aditya Challapally. finance.yahoo.com/news/mit-report-95-generative-ai
  4. Legal.io. (2025, August 23). MIT Report Finds 95% of AI Pilots Fail to Deliver ROI, Exposing "GenAI Divide". legal.io/blog/MIT-Report-Finds-95-of-AI-Pilots-Fail
  5. Snyder, J. (2025, August 26). MIT Finds 95% Of GenAI Pilots Fail Because Companies Avoid Friction. Forbes. forbes.com/sites/jasonsnyder/2025/08/26/mit-finds-95-of-genai-pilots-fail
  6. Pravitech. The GenAI Divide: Why 95% of Enterprise AI Pilots Are Failing. pravitech.substack.com/p/the-genai-divide
  7. Congruity 360. (2025, October 28). Why 95% of Generative AI Pilots Are Failing, And How Data Quality Is the Missing Link. congruity360.com/blog/why-95-of-generative-ai-pilots-are-failing
  8. Legal.io. (2025). Reporting sample figures of 52 executive interviews and 153 leader surveys, differing from other secondary accounts of the same study.
  9. Trullion. (2025, September 8). Why 95% of GenAI Projects Fail, And Why The 5% That Survive Matter, citing Gartner findings on agentic AI project stall rates. trullion.com/blog/why-95-of-ai-projects-fail
  10. Dataiku. (2025, October 16). MIT Says 95% of GenAI Pilots Fail: Here's How To Beat The Odds, containing the 18-month window projection. dataiku.com/blog/moving-past-genai-pilots
  11. Gallacher, D. Beyond ROI: Are We Using the Wrong Metric in Measuring AI Success? A Response to MIT's "95% AI Failure" Study. UC Berkeley Executive Education. exec-ed.berkeley.edu/.../Beyond-ROI-pdf.pdf

This article discusses published research and commentary and is provided for general informational purposes. It is not technology procurement or investment advice. Technology adoption decisions should be evaluated against your own business's specific circumstances with appropriate professional input.