In 2020 BCG studied 70 digital transformations and surveyed 825 senior executives. 30% hit their targets. 44% created some value and fell short. 26% created almost nothing.
In March 2025 S&P Global Market Intelligence surveyed more than 1,000 companies across North America and Europe. 42% had scrapped most of their AI initiatives that year, up from 17% the year before. On average, 46% of proofs of concept were killed before they reached production.
Six years apart. Two different technologies. The same arithmetic.
The thesis
Your AI transformation will fail where your digital transformation failed: at the handoff between the people who understand the business and the people who build. That part is not new. What is new is the artefact that has to cross the handoff, and the fact that nobody has been assigned to own it.
In digital transformation the artefact was requirements. Requirements have a named owner, a template, a sign-off and a change process. Every organisation over 200 people has a settled opinion about who writes them.
In AI the artefact is the evaluation set. It is a dated, versioned list of real inputs, each paired with a judgement on whether the output was correct, written by somebody who knows the business. Almost nobody owns it. The business side assumes it is technical because it lives next to the code. Engineering assumes it is a business input because an engineer cannot say whether a claims summary is right. So it never gets written, and the system ships against taste.
What actually changed
I have written software since 1995. For most of that time the hard part of a handoff was ambiguity in a spec. You wrote the requirement, an engineer misread it, a test caught the gap. The test was cheap because correct was a property of the code. Two plus two returns four or it does not.
A language model moved correct out of the code and into the business. Whether a summary of a 60 page annual report is right is not a property of the function. It is a property of what your analysts would have written, and the only people who can say are your analysts.
That is a job. It has a size. Hamel Husain, who teaches this discipline to engineering teams, puts it at 60 to 80% of development time spent on error analysis and evaluation. His method is specific: appoint one principal domain expert as the final word on quality, have that person personally annotate at least 30 traces before anything is automated, push to 100 or more, and make every judgement binary pass or fail rather than one to five, because a five point scale lets an annotator hedge into the middle.
Read that as an org chart problem rather than an engineering one. Between 60 and 80% of the work on your AI system belongs to a person whose day job is claims, or lending, or clinical coding. Nobody has taken that person off their normal duties. Nobody has written the role into the program plan. The steering pack reports adoption, licences and pilot counts, because those are countable without a domain expert. So the build runs for nine months with no measured definition of correct, and then somebody asks what it returned.
Three places this shows up
Buy beats build by exactly this margin. MIT NANDA reviewed 300 plus public AI deployments between January and June 2025, with 52 structured interviews and 153 survey responses. 95% of enterprise GenAI pilots produced zero measurable return against 30 to 40 billion dollars of spend. Tools purchased from specialist vendors reached deployment about two thirds of the time. Internal builds managed about one third. The usual reading is that vendors are better engineers. I do not accept that. A vendor arrives carrying an evaluation set already built across dozens of customers, and it encodes what correct looks like for that task. Your internal team starts with an empty file and no authority to requisition the one person who could fill it.
The same instrument gap repeats at the cost layer. KPMG surveyed 204 US C-suite leaders at companies above 1 billion dollars in revenue between 28 April and 25 May 2026. 53% have deployed agents. 18% orchestrate multiple agents, double the previous quarter. Average planned investment is 202 million dollars over the next twelve months. Only 26% have real-time visibility of what those agents cost to run, and only 36% have token or usage controls. In asset management and private equity, full cost visibility sits at 4%. Same shape as the evaluation gap: money committed at scale, no instrument attached, and the numbers came from the people who signed the cheque.
Gartner’s forecast lands on top of that. More than 40% of agentic AI projects cancelled by the end of 2027, on escalating costs, unclear business value and inadequate risk controls. Unclear business value is not a cause. It is the name we give a missing measurement once the invoice arrives.
I got the order wrong myself. D30 pulls three statement fact files out of ASX annual reports. I built the extraction engine before I wrote the articulation gates, which is backwards, and I paid for it in a rebuild. The gates are now the definition of correct and the model does not get a vote on them. SearchFIT’s unit is one tracked prompt scanned across the answer engines, and whether the answer is right is a judgement about brand positioning that a marketing lead has to make, not an engineer. D23 runs managed Apache Superset, where a dashboard is correct when the finance lead would put their name on it. In all three the definition of correct has one person against it. In client programs I now decline to start a build until that name exists.
Where this breaks
Not every task has a correct output. Ideation, drafting, exploratory analysis: there is no single right answer, and forcing a pass or fail set produces a metric that punishes range. Preference comparison between two candidate outputs is the right instrument there, and it is a different build with a different cost.
Gartner puts the constraint upstream of mine. Its February 2025 release, drawn from 1,203 data management leaders surveyed in July 2024, predicts organisations will abandon 60% of AI projects through 2026 for want of AI-ready data, and reports 63% either lacking the right practices or unsure whether they have them. If you cannot assemble the data you never reach the handoff, and the evaluation set is moot. My answer is that these are the same failure at two depths. That is an argument, not a proof.
Evaluation sets decay. A set written against one model version can be passed by a newer model that is worse at the thing you actually care about, because the set encodes last year’s failure modes. That is why 60 to 80% is ongoing spend and not a setup cost. Any executive who signs it as a one-off has signed the wrong thing.
The 70% transformation failure statistic also deserves the scepticism it gets. BCG’s study is 70 companies and 825 executives, not a law of nature. I use it because the S&P and MIT numbers were collected differently, five years later, and landed in the same band.
What I would do on Monday
Name one person per AI system as the principal domain expert. One, not a committee. Put the name in the program plan next to the engineering lead.
Give that person 20% of their week back. If you cannot fund one day a week out of an operational team, you cannot fund the project either.
Get 30 real production traces annotated by that person this week. Binary, pass or fail, one line of reason on every fail. Reach 100 within a month.
Move the evaluation set into the same repository as the code, dated and versioned, and report its pass rate to the steering committee instead of adoption or licence counts.
Cancel any pilot that cannot state its correct output in one sentence. It gets cancelled in 2027 anyway, and it costs more then.
The number that did not move
30% in 2020. 42% in 2025. The failure rate held because the failure held. All that changed was the name of the document nobody would own.
Write the definition of correct. Put a name against it. The rest is engineering.
I run AI transformation programs through PADISO. If you are trying to work out what actually changes in your operating model, book a call.

