Why 95% of Enterprise AI Projects Fail: The Scoping Discipline That Beats the Odds
MIT found 95% of integrated enterprise AI pilots delivered no measurable P&L impact. The models were rarely the reason. What separated the 5% was how they decided what to build, in what order, and with which tool.
In the summer of 2025, MIT’s NANDA initiative published The GenAI Divide: State of AI in Business 2025, and one number escaped the report and never came back: 95% of organizations getting zero return on generative AI, despite $30 to $40 billion in enterprise spending. Only 5% of integrated pilots were producing measurable P&L impact. Gartner’s forecast of more than 40% of agentic AI projects canceled by the end of 2027 points the same way.
The models were not the reason. Foundation models, agent frameworks and vector databases are more capable than they were a year ago, and the MIT report says so itself: the divide “does not seem to be driven by model quality or regulation, but seems to be determined by approach.” The approach that fails is how enterprises decide what to build, in what order, and with which tool. On episode 93, Jacob Andra, CEO of the AI consulting firm Talbot West, named the pattern:
“The most common issue we see is a client who has already decided the answer is an LLM, and they want us to help them wrap a workflow around it. That is scoping in reverse.”
Jacob Andra, CEO of Talbot West, on episode 93
What did MIT’s “95% of AI pilots fail” study actually measure?
A statistic this popular deserves a look at its ingredients.
The report, dated July 2025, drew on structured interviews at 52 organizations, a survey of 153 senior leaders and a review of more than 300 publicly disclosed AI initiatives between January and June 2025. MIT’s own copy sits behind a form, so the version most people read is a mirrored PDF of the GenAI Divide report. Its definition of success is specific: “deployment beyond pilot phase with measurable KPIs. ROI impact measured 6 months post-pilot, adjusted for department size.” So the 95% describes integrated pilots that had not produced measurable profit-and-loss impact roughly six months on. It does not mean the models broke, and it does not mean the projects were cancelled. The authors add that the figures are “directionally accurate based on individual interviews rather than official company reporting,” and that six months “may be insufficient to fully assess ‘successful deployment’ for complex enterprise systems.” Treat 95% as a pattern, not a census.
The pattern is still striking, and two details inside it matter more than the headline. First, the funnel. Generic chatbots did fine: 80% of enterprises investigated them, 50% piloted, 40% implemented. Custom, task-specific tools, the ones wired into a real workflow, were evaluated by 60%, piloted by 20%, and reached production at 5%. Second, the people. The report found 90% of employees using LLMs regularly while only 40% of companies had bought a subscription, the shadow economy we wrote about in Shadow AI: the risks, and how to govern it. Individuals were getting value every day; the enterprise pilot pipeline was where it died. That is a scoping problem before it is a capability problem.
Two other studies keep the MIT number honest company. RAND’s 2024 report on the root causes of AI project failure, built on interviews with 65 data scientists and engineers, opens with “by some estimates, more than 80 percent of AI projects fail—twice the rate of failure for information technology projects that do not involve AI.” Gartner predicted in 2024 that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025, and in 2025 that over 40% of agentic AI projects will be canceled by the end of 2027 for “escalating costs, unclear business value or inadequate risk controls.”
Three studies, three definitions, one shared diagnosis: every cause on those lists is a decision made before the first line of code. Which problem, which data, which tool, which controls. Read honestly, the headline says most enterprises have not yet learned how to choose and sequence AI work. That is a fixable problem, and fixable problems are good news.
Why do enterprise AI projects fail? Scoping in reverse
MIT’s data shows how the 5% differed from the rest, and the difference is procedural. Pilots built with outside partners and purchased, customized tools reached deployment about 67% of the time; internally built tools, about 33%. Our read of why: outside teams ask harder questions earlier, and they insist on a measurable baseline before building anything, because their fee depends on the outcome being real. Talbot West’s own site describes the firm as “a multidisciplinary modernization practice” that “resell[s] nothing” and takes “no platform commissions,” which is worth checking about any advisor: one with no platform to sell has no reason to pick the tool first.
The report’s explanation of the stalls reads like a scoping post-mortem: “Most fail due to brittle workflows, lack of contextual learning, and misalignment with day-to-day operations.” It also found that sales and marketing captured about 70% of AI budgets while less glamorous back-office functions delivered the better returns. Money followed visibility rather than value.
Three scoping mistakes repeat across the failed projects we hear about, and they map cleanly onto RAND’s root causes:
- Solving a problem the business is not actually paying to solve. RAND: “industry stakeholders often misunderstand—or miscommunicate—what problem needs to be solved using AI.” The pilot works and nobody’s number moves, because no number was ever named.
- Ignoring the data pipeline that feeds the model. RAND: “many AI projects fail because the organization lacks the necessary data to adequately train an effective AI model.” The AI performs on a clean sample and fails on production data.
- Treating AI as a single capability instead of a toolkit. RAND: the organization “focuses more on using the latest and greatest technology than on solving real problems for its intended users.” Committing to an LLM before asking whether an LLM is the right fit.
Andra’s “scoping in reverse” is all three at once: the tool arrives first, then a problem is recruited to justify it. The discipline that beats the odds runs the other way, and it starts with a map.
How to map AI dependencies before you pick a tool
Before selecting a model or a vendor, draw the full dependency map of the workflow the AI will touch, from the trigger that produces the input to the business number that is supposed to change. Talbot West’s FRAME methodology, per the firm’s own description, “maps complex operations and sequence change” and “produces a sequenced roadmap rather than a list of recommendations.” Whatever you call yours, it has four layers.
Data. Where does the input live, and in which system of record? How clean is it in production, as opposed to the sample the demo ran on? Who owns it, and what classification does it carry? An AI data bill of materials is the formal version of this layer: every dataset the system touches and where each one came from.
Process. What happens before the AI call, and what happens after it? What breaks downstream if the output format changes by one field? What is the baseline today, in the unit the business cares about, so the pilot has something to beat?
Human. Who reviews the output, and how often? Who can override it, and how? Who is accountable when it is wrong? If those three names are the same person, or nobody, the map has found its first problem.
Security and compliance. Which data classifications does the call touch? What audit trail and retention does the regulator, the customer contract or your own SOC 2 scope require? Whose approval is needed: security, legal, the budget owner? The cost of this layer is often larger than the model cost itself, especially in regulated sectors. It is also where a security team gets to be the enabler, the posture we argued for in Stop Saying No: give security the map early and the review becomes a lookup rather than a late veto. If the AI is an agent with its own credentials, an AI access control audit turns the access questions into a list with owners.
The output of this step is not a tool choice. It is a scoped problem statement: this workflow, this input, this number, these people, these controls. If you are evaluating a vendor’s agent rather than building your own, the same four layers are what an AI readiness assessment checks before the contract is signed, and the declaration a vendor should hand you up front is the subject of AI agent governance starts at onboarding.
Finding the iceberg under the demo
Andra has a name for the mapping exercise: “finding the iceberg under the demo.” The demo is the tip: one model, one prompt, a clean sample, the happy path. It is cheap, fast and genuinely impressive, which is why leadership asks how soon it can ship.
Below the waterline sit the dependencies in the figure: the production data pipeline, the downstream format contracts, the review and override roles, the audit trail, the approvals, the training, the feedback loop. None of it is exotic, and all of it is invisible from the demo. MIT’s description of the 5% is a description of people who looked below the line first: they “demand process-specific customization and evaluate tools based on business outcomes rather than software benchmarks.”
When is an LLM the wrong answer? Pick the approach after the map
Only after the dependency map is complete do you pick the AI approach, and this is where the 95% go wrong most often: assuming the answer is an LLM because that is what leadership has heard of.
“LLMs are one tool in a toolbox. They are extraordinary at some things. They are genuinely bad at others.”
Jacob Andra, Talbot West, episode 93
Four categories of work usually do better without an LLM as the primary engine. Deterministic decisions with known rules: a rules engine beats an LLM on accuracy and on auditability, which matters the moment a regulator asks why. Numeric forecasting on structured data, where classical models have decades of head start. Structured extraction into a fixed schema. And anything that requires a numeric guarantee, because a model that is right most of the time is wrong on some invoice every week.
RAND’s fifth root cause is the same warning from the research side: “AI projects fail because the technology is applied to problems that are too difficult for AI to solve,” and its matching recommendation is simply to “understand AI’s limitations.” The right pattern is often a blend: an LLM front end for language, classical models for the numeric work, a symbolic layer for the decisions that must be correct and explainable. Stephen Karafiath, Talbot West’s co-founder and SVP of AI Solutions Architecture, joined Andra for that part of the conversation, and the architecture they described has its own post in Neuro-Symbolic AI: Why Enterprises Need More Than Large Language Models. For scoping purposes the rule is shorter: the map tells you which parts of the workflow must be correct, and those parts do not get a probabilistic engine.
Why the human layer decides whether a pilot becomes a business system
Even a technically perfect deployment stalls if the people downstream do not adopt it. Three patterns recur. Nobody is trained to use the output, so usage never climbs past the people who sat in the pilot. Nobody is trained to override it, so staff either rubber-stamp everything or reject everything. And nobody owns the feedback loop, so the model drifts and the first person to notice is a customer.
“If your pilot succeeds but your people do not adopt it, you have built a science project, not a business system.”
Jacob Andra, Talbot West, episode 93
MIT arrived at the same layer from the data. “The core barrier to scaling is not infrastructure, regulation, or talent. It is learning,” the report says: most systems “do not retain feedback, adapt to context, or improve over time,” and people abandon tools that cannot remember last week. In most workflows the learning happens through a human who corrects the output and a loop that carries the correction back. Change management is a line item, not a footnote.
Ged Ossman, founder and CEO of Interf, whose company works on getting AI agents onboarded into enterprises, sees the same thing from the vendor side. On episode 99 he argued that the technology was never the hard part:
“The core problem is not even slow adoption. It was always about communicating, aligning stakeholders internally on how we can adopt this solution in a way that does not introduce new risk.”
Ged Ossman, founder and CEO of Interf, on episode 99
That internal process, he said, “is the hardest part.” It is also the part a dependency map makes tractable, because the human and approvals questions get names attached in week one. Security tooling learned this lesson the expensive way: shift-left programs kept failing because nobody redesigned the developer’s workflow around the scanners. AI pilots stall the same way, and the fix is the same: design the human workflow as carefully as the model call.
How to scope an AI project: Talbot West’s three-round framework
Talbot West runs a three-round process, and it is compact enough to borrow whole.
Round 1, problem cataloging. List every workflow where AI could plausibly help. Score each on business value and on data readiness. Pick no tool yet. The deliverable is a ranked list, and the discipline is resisting the one candidate that already has a demo attached.
Round 2, dependency and fit analysis. For the top candidates, draw the four-layer map above and identify the right AI class for each: an LLM, a classical model, a rules engine, a blend, or the honest option, “not AI.” The deliverable is a scoped problem statement per candidate, plus the dependencies that must exist before a pilot can be honest.
Round 3, pilot design with a measurable baseline. Design the narrowest pilot that can prove the value. Set the success threshold and the time box before it starts, and stop it on schedule if it misses. MIT measured its 5% at six months post-pilot and noted that even that may be short for complex systems; your threshold and your date belong in writing before the kickoff. A pilot that stops on time with a clear miss is a scoping success, because the cheapest failed project is a short one.
The structure costs more up front than “pick a tool and prototype.” What it buys is that every dependency in the iceberg gets met on purpose, in a meeting, instead of by surprise, in production. Gartner’s three cancellation reasons, escalating costs, unclear business value and inadequate risk controls, are all settled in round two, before any money is spent on round three.
What the 95% means for a startup selling AI into enterprises
Most of our readers are on the other side of this table: founders and CTOs selling an AI product to the enterprises MIT surveyed. Everything above is a description of your buyer, and your buyer’s dependency map has a box with your name on it.
Every layer of that map has a question about you. The same buyers carry ten reasonable objections to agentic workflows into every evaluation, and a checklist of tells for vendors who overclaim; we catalogued both, with the evidence on each, in The Art and Zen of Adopting AI. Data: which of their records does your model touch, and does any of it leave their tenant? Process: what happens downstream when your output format changes in a release? Human: who on their side can override your output, and what do you show them so they can? Security and compliance: what classification is involved, what audit trail do you produce, what do their reviewers have to sign? MIT’s finding that externally sourced tools reached deployment twice as often as internal builds is your opening, with one condition: the buyers who succeeded “evaluate tools based on business outcomes rather than software benchmarks.” Show up with their number, not your leaderboard.
The vendors who clear the pilot answer the layers before being asked, which in practice means three artifacts. A dependency declaration: the data, systems, identities and approvals your product needs to work, which Ged’s onboarding protocol formalizes and which you can write in a document today. A data bill of materials for your own model, so “what was this trained on” has an answer; if customer data is anywhere near your training or evaluation sets, PII and secret scrubbing is the fixable with a known price. And compliance evidence that matches the buyer’s map: a SOC 2 report that names your AI components in its system description, and, increasingly, ISO 42001, the AI management system standard that turns “how do you govern this model” into a certificate instead of a forty-question thread.
Who does that work at a startup is the same question as who does any security work: in-house, on-demand, or YOLO. The right answer depends on your stage, and it is the subject of our live founder webinar on October 6.
The scoping questions to ask before your next AI project
Whether you are buying, building or selling, six questions separate the 5% from the rest, and all six can be answered before a dollar goes to a model.
- Which business number is this supposed to move, and what is it today?
- Where does the input live, who owns it, and how clean is it in production?
- What runs before and after the AI call, and what breaks if the output changes?
- Who reviews, who overrides, and who is accountable when it is wrong?
- Which data classifications, audit trails and approvals are involved?
- Given all of that, which class of AI fits, and does any of it need to be an LLM at all?
Answer them and you have a scoped problem statement, and the odds have already moved in your favor. If the security and compliance layer is the one nobody on your team owns, our ISO 42001 practice builds the AI management system that makes the map a repeatable control, and the evidence your enterprise buyers are starting to ask for. If you would rather hear the argument first, Jacob Andra and Stephen Karafiath make it in 45 minutes on episode 93.
AI project scoping frequently asked questions
- What percentage of AI projects fail?
- It depends on what is measured. MIT NANDA's 2025 GenAI Divide report found 95% of organizations getting zero return from generative AI pilots, defined as no measurable P&L impact six months after the pilot. RAND cites estimates that more than 80% of AI projects fail, twice the rate of ordinary IT projects. Gartner predicted at least 30% of generative AI projects abandoned after proof of concept by end of 2025, and over 40% of agentic AI projects canceled by end of 2027. The common thread is scoping, not model quality.
- What did the MIT study on 95% of AI pilots actually measure?
- Whether integrated generative AI pilots produced measurable P&L impact. The report defined success as deployment beyond the pilot phase with measurable KPIs, with ROI measured six months post-pilot, and found only 5% of pilots met that bar. The data came from interviews at 52 organizations, a survey of 153 senior leaders and a review of 300 public deployments between January and June 2025. The authors call the figures directionally accurate rather than official company reporting, so treat 95% as a pattern, not a census.
- Why do most enterprise AI projects fail?
- Because the tool was chosen before the problem was scoped. Three mistakes repeat: solving a problem the business is not paying to solve, ignoring the data pipeline so the model works on a clean sample and fails on production data, and treating AI as one capability instead of a toolkit. RAND's interviews found the same pattern: stakeholders misunderstand the problem, lack the data, or focus on the latest technology instead of the user's problem. All of it is visible in advance if someone maps the dependencies first.
- What is a dependency map for an AI project?
- A diagram of everything the AI call touches, drawn before any model or vendor is selected. It has four layers: data (where the input lives, how clean it is, who owns it), process (what runs before and after the call, what breaks if the output changes), human (who reviews, overrides and is accountable) and security and compliance (data classification, audit trail, approvals). Its output is a scoped problem statement rather than a tool choice. Jacob Andra of Talbot West calls the exercise finding the iceberg under the demo.
- How do you scope an AI project before choosing a tool?
- Run three rounds. First, catalog every workflow where AI could plausibly help and score each on business value and data readiness, choosing no tool yet. Second, map the full dependency graph for the top candidates and identify the right AI class for each, including the option of no AI. Third, design the narrowest useful pilot with a measurable baseline, a success threshold and a time box set before it starts, and stop it if it misses. The tool choice happens after round two, never on day one.
- When is a large language model the wrong tool for an enterprise AI project?
- When the output must be exactly right, reproducible or numerically guaranteed. Deterministic decisions with known rules belong in a rules engine, which beats an LLM on accuracy and auditability. Numeric forecasting on structured data belongs with classical models. Structured extraction into a fixed schema and anything needing a numeric guarantee also usually do better without an LLM as the primary engine. LLMs excel at language, so the common winning pattern is an LLM front end with deterministic or classical components doing the work that has to be correct.
- How long should an AI pilot run before you decide?
- Set the answer before the pilot starts. Agree the baseline metric, the success threshold and the end date up front, then hold to them; a pilot that can be extended indefinitely is a science project with a budget line. For reference, MIT's study measured ROI six months after the pilot, and the authors note that even six months may be short for complex enterprise systems. A time-boxed pilot that misses its threshold and stops on schedule is a successful scoping outcome, because the cheapest failed project is a short one.
- Does AI scoping discipline matter for a startup selling AI to enterprises?
- Yes, because your product sits inside your buyer's dependency map. Their data, process, human and security layers all have questions about you: what data your model touches, what happens when your output format changes, who overrides it, what audit trail you provide. Vendors who answer those before the pilot reach deployment more often; MIT found externally sourced tools reached deployment about twice as often as internal builds. Pairing a clear dependency declaration with SOC 2 or ISO 42001 evidence turns the security review into a lookup instead of a delay.