Zum Inhalt springen
AI

Why Most AI Projects Never Pay Off (And How to Be in the Minority)

Most enterprise AI pilots quietly die without earning back their cost. The real reasons they fail — and the discipline that separates the ones that don't.

Aparajita Singh
Aparajita Singh
Veröffentlicht
Lesezeit6 min
Why Most AI Projects Never Pay Off (And How to Be in the Minority)

Nobody announces the failure. The pilot runs, the demo goes well, a report gets written, and the project quietly stops being mentioned. Twelve months later the budget has moved and nobody can say precisely what happened.

That pattern repeats across enough organisations to be worth taking seriously, and the causes are consistent enough to be avoidable.

The projects start from the wrong end

The most common failure is structural and happens before any code is written: the project starts with the technology rather than the problem.

Someone sees a capability, imagines an application, and commissions a pilot. The pilot is technically successful — the model does the thing — and then nobody can work out how to put it into production, because it was never attached to a process anyone owns or a cost anyone was trying to reduce.

The inverse works. Start with a specific, expensive, repetitive process. Establish what it currently costs in hours and errors. Then ask whether this technology reduces that number. Projects framed this way tend to survive, because there is a person whose job gets better and a figure that moves.

A blunt test: if you cannot name who will use it daily and what they currently do instead, the project is a demo with a budget.

Pilot success is not evidence

The second failure is mistaking a good pilot for a working system.

Pilots run on curated data, with the team's attention, on the cases someone chose. Production runs on whatever arrives, unattended, including the malformed, the ambiguous, and the actively adversarial. The gap between those two is where most projects die, and it is not a modelling gap.

A useful pilot design is deliberately unflattering: run it on a random sample of real inputs including the messy ones, unattended, and measure how often it needs a human. That number is your actual operating cost, and it is usually higher than the demo suggested.

Two paths from pilot: one measures curated inputs with the team watching and reports high accuracy; the other measures random real inputs unattended and reports an intervention rate — only the second predicts production cost

Fig. — A pilot that cannot fail is not measuring anything.

The costs that appear after the demo

Business cases routinely count model or licence cost and stop there. The costs that actually accumulate:

Integration. Connecting to real systems, with real authentication and real error handling, is usually the largest line and rarely the estimated one.

Data preparation. Almost every project discovers its data is worse than assumed — duplicated, inconsistent, missing the field the whole approach depended on. This is not a detour; for many projects it is the majority of the work.

Human review. Anything with consequences needs someone checking a proportion of outputs. That is an ongoing role, not a launch task, and it is the cost most likely to be omitted entirely.

Maintenance. Inputs drift, upstream systems change, models get deprecated. A deployed system needs someone who owns it, and the day nobody does is the day it starts quietly degrading.

The projects that pay off tend to be the ones where these were priced honestly and the answer was still yes.

Adoption is where the value is decided

A technically excellent system that people work around returns nothing, and this is more common than technical failure.

The patterns are predictable. Staff were not consulted and assume it threatens their jobs. The tool adds a step rather than removing one. Nobody explained what it is bad at, so the first confident mistake destroys trust permanently. There is no obvious way to correct it when it is wrong, so people stop trying.

The fix is unglamorous: involve the people doing the work while designing, be specific about limitations, make correction one click, and measure adoption rather than assuming it. A weekly number showing how often the tool is used versus bypassed tells you more than any accuracy metric.

Build, buy, or wait

A quieter cause of wasted budget: building something that arrived in a product six months later.

The capability landscape moves fast enough that a custom system solving a general problem — summarising documents, drafting replies, classifying tickets — often gets overtaken by a feature in software you already pay for. Teams that spent two quarters building it end up maintaining a worse version of something now included in their licence.

The distinction worth drawing: build where the value comes from your own data, your own process, or your own domain knowledge. Buy where the problem is general. Wait where the problem is general and nobody has shipped it well yet, because they will.

That last option is legitimate and rarely chosen, because "we decided not to build this yet" makes a poor slide. It is frequently the highest-return decision available.

What the successful ones have in common

They are narrower than expected. One process, one team, a clear definition of done. Broad transformation programmes have a much worse record than boring departmental automations.

They have a named owner with authority to change the process, not just the software. Most of the value in automating a workflow comes from redesigning it, and that needs someone who can.

They measure against a baseline that existed before the project. Teams that did not record what the process cost beforehand cannot demonstrate improvement afterwards, which is how successful projects still get cancelled.

And they treat the first version as deliberately limited — assisting rather than deciding, with a human in the loop — then widen autonomy as evidence accumulates. The ones that launched fully autonomous almost all pulled back after an incident, at a cost to credibility that outlasted the technical fix.

The uncomfortable question worth asking early

Before committing budget: if this works exactly as intended, what specifically changes?

If the answer is a number — hours saved, errors avoided, a queue cleared, revenue captured that currently leaks — you have a project. If the answer is a capability, an efficiency, or a modernisation, you have an aspiration, and aspirations are where AI budgets go to disappear.

There is a related discipline worth adopting: decide in advance what would make you stop. A pilot with no kill criteria runs until someone loses patience, and by then the sunk cost makes cancelling politically expensive. Agreeing beforehand that an intervention rate above some threshold means abandoning the approach turns a failure into a cheap, early answer rather than a slow, awkward one.

The organisations getting real returns are rarely the ones with the most ambitious programmes. They are the ones who picked something tedious, measured it, and finished.

Aparajita Singh
Geschrieben von

Aparajita Singh

Haben Sie ein Projekt im Sinn?

Erzählen Sie uns davon — wir antworten innerhalb eines Werktags mit einer ehrlichen Einschätzung zu Fit und Umfang.