Programmes rarely fail in novel ways. They fail in a small number of recognisable patterns, most of which were determined before implementation began, and most of which were visible for months before anyone said so.
That is the useful framing. "Why do transformations fail?" is an unanswerable question that produces platitudes about change management. "Which failure is currently underway?" is answerable, and has a different intervention attached to each answer.
One caveat, stated plainly: this taxonomy is an argued framework built from our decision model and its conditions, not a statistical study of programme outcomes. It should be judged on whether it is diagnostically useful, not on a sample size.
The nine modes
| # | Failure mode | Decided at | Early warning | Still works once underway |
|---|---|---|---|---|
| 1 | Automating chaos | Before selection | Requirements gathering finds three ways to do everything | Narrow scope to the stable subset |
| 2 | Bolting onto a capable ERP | Before selection | Nobody can say what the ERP module actually lacks | Pause; run the gap properly |
| 3 | Buying against the wrong ecosystem | At selection | Integration is "on the roadmap" | Renegotiate scope; budget the integration honestly |
| 4 | No process owner with authority | At mobilisation | Design decisions escalate and return unresolved | Appoint one, publicly, with decision rights |
| 5 | Scope defined by modules, not outcomes | At mobilisation | The plan is a list of features with dates | Re-anchor each workstream to a measurable outcome |
| 6 | Ignoring a binding constraint | At mobilisation | The plan assumes hiring or spend nobody has approved | Re-plan inside the constraint; record the deferred decision |
| 7 | The pilot that cannot graduate | Mid-delivery | Pilot succeeds; production date keeps moving | Test control and capability explicitly, not the technology |
| 8 | Benefit realisation with no baseline | At business case | Nobody can state today's number precisely | Baseline now; accept the gap in the record |
| 9 | Adoption assumed, not designed | Throughout | Training is a line item in the final month | Redesign the work, not the training |
Modes 1, 2, 3, 6 and 8 are all decided before or at selection. That is the framework's central observation: the majority of finance transformation failure is a decision-quality problem wearing an execution costume. Programmes get blamed on delivery because delivery is where the symptoms appear.
The pre-selection failures
1 · Automating chaos
Encoding an unstandardized process into software. The implementation either forces a single path — process design, now happening mid-build, on day rates, by people without authority to decide — or encodes every variant, making the variation permanent and harder to change than it was in a spreadsheet.
Warning: requirements workshops keep discovering that different entities or preparers do the same task differently, and each difference is treated as a configuration requirement rather than a decision to be made.
Intervention: narrow scope to the processes that genuinely meet the standardization test. A smaller programme that works buys the credibility to fund the rest.
2 · Bolting onto a capable ERP
Buying a third-party platform to do something the existing ERP already does, usually because the module was never implemented or was implemented badly and abandoned.
Warning: ask what specifically the ERP module cannot do. If the answer is a feeling rather than a functional gap, this mode is active.
Intervention: pause and run the gap analysis properly. This is uncomfortable mid-programme and cheaper than the alternative — which is a second platform, a permanent integration, and a reconciliation obligation between two systems of record.
3 · Buying against the wrong ecosystem
Selecting on capability while under-weighting integration to the ERP you actually run. Integration is where finance implementations overrun, and a generic connector transfers that risk from the vendor to you.
Warning: the integration is described as "supported" but the reference customers run a different ERP, or the native path is "on the roadmap."
Intervention: renegotiate scope so integration is the vendor's obligation with a defined acceptance test, and budget it at what it will actually cost.
8 · Benefit realisation with no baseline
The business case promises days, headcount or cost, and nobody recorded today's number precisely enough to prove it later. Twelve months on, the programme cannot demonstrate value even where value was delivered — which poisons the funding of the next programme.
Warning: the current-state number in the business case is round, or sourced from "management estimate."
Intervention: baseline now, and record in the programme log that the baseline post-dates the start. An honest late baseline is worth considerably more than a reconstructed one.
The mobilisation failures
4 · No process owner with authority
Standardization is a decision-making exercise, not a documentation one. Someone must be able to say "we will all now do it this way" and make it stick across entities that do not report to them.
Warning: design decisions escalate to a steering committee and come back unresolved or softened into optionality.
Intervention: appoint one owner, publicly, with explicit decision rights and an escalation path that terminates. The absence of this role is the most common cause of mode 1 persisting after it has been identified.
5 · Scope defined by modules, not outcomes
The plan is a list of things to be switched on, each with a date. It can be fully delivered without anything getting measurably better — and frequently is.
Warning: the programme reports percentage completion of features rather than movement in an outcome measure.
Intervention: re-anchor each workstream to a measurable outcome with a baseline. Any workstream that cannot be tied to one should be challenged on whether it belongs.
6 · Ignoring a binding constraint
A budget freeze, a headcount cap, an ERP migration already in flight. The plan assumes spend or hiring that has not been approved, and quietly slips as reality arrives.
Warning: the resourcing plan depends on recruitment nobody has authorised, or on capacity from a team already committed elsewhere.
Intervention: re-plan inside the constraint and record separately what would be correct without it. That record is what makes the decision cheap to revisit when the constraint lifts — and it is almost never kept.
The delivery failures
7 · The pilot that cannot graduate
The pilot succeeds. The production date moves, repeatedly, without anyone saying the programme has stopped. This is the characteristic failure of finance AI work specifically.
The cause is that a pilot tests the technology, while production requires control and capability — evidence of what the model did, review where materiality demands it, and people who can tell whether an output is wrong. A pilot that does not deliberately test those two dimensions has not de-risked the thing that will block it.
Warning: the pilot's success criteria are all technical, and the phrase "we just need to industrialise it" appears.
Intervention: stop extending the pilot. Test control and capability explicitly, as their own workstream, with the audit function in the room.
9 · Adoption assumed, not designed
The system works, and people work around it. Usually because the new process is genuinely worse for the person doing it — more clicks, less discretion, no visible benefit to them.
Warning: training appears as a single line item in the final month, and no one has described the target-state role of any individual.
Intervention: redesign the work rather than the training. If the new process is worse for the person performing it, no amount of communication fixes that.
Running the diagnostic
The framework is only useful applied to a live programme, which is politically harder than applying it to a past one.
- Diagnose the mode before proposing the fix. Interventions differ sharply; the wrong
one applied confidently makes things worse and burns the diagnosis.
- Ask when the failure was decided, not who decided it. Modes 1, 2, 3, 6 and 8 were
determined before delivery — which means the delivery team is usually not the problem, and saying so is what makes the conversation possible.
- Expect more than one. Modes 1 and 4 travel together almost always; 5 and 8 do too.
- Score against the early warnings, not the outcomes. By the time the outcome is visible,
the cheap interventions have expired.
The strongest objection is that a failure taxonomy is retrospectively convenient — every programme exhibits several of these patterns to some degree, including successful ones.
That is a fair charge against most such lists, and it is the reason to be careful with this one. Nine categories broad enough to cover finance transformation will match any programme looked at hard enough, which makes the framework feel more predictive than it is. Confirmation is easy; falsification is not offered.
Two things narrow the exposure. The framework is built to be diagnostically discriminating rather than descriptively complete — each mode carries a distinct intervention, so misdiagnosis has a visible cost and the categories are not interchangeable. And each mode has a stated early warning that is observable before the outcome, which is the only form in which this kind of framework can be wrong in a useful way.
What it genuinely cannot do is tell you which mode dominates. A programme showing signs of four modes needs judgement about which is load-bearing, and this framework does not supply that judgement — it only ensures the candidates are named. Anyone claiming a failure taxonomy does more than that is overselling it.
- Most finance transformation failure is decided before implementation begins. Five of
the nine modes are set at or before selection — which is why blaming delivery usually misdiagnoses the problem.
- Diagnose the mode before applying a fix. The interventions differ sharply, and the
cheap ones expire well before the outcome becomes visible.
- The pilot that cannot graduate is the signature finance AI failure, and it happens
because pilots test technology while production requires control and capability.
Where this leaves you. Take the programme currently in flight and score it against the nine early warnings rather than the nine outcomes. Anything scoring on a pre-selection mode needs a decision revisited, not a delivery push.