The question arrives from the board, and it arrives without a definition. Are we ready for AI? Finance leaders answer it in one of two unsatisfactory ways: with enthusiasm that has no basis, or with caution that has no argument. Both are guesses. Neither survives a follow-up question.
The problem is that readiness is being treated as a single property, when it is five separate ones that fail independently. An organisation can have excellent data and no control framework. It can have a mature control framework and no process stability. Each of these produces a different failure, at a different time, with a different cost.
This framework makes the question answerable. Five dimensions, each scored on the same 1–5 spine we use across all our research, each with a defined action at each level.
The governing rule
Readiness is governed by the weakest dimension, not the average. The dimensions are not substitutable: excellent data does not compensate for absent controls, and a strong control framework does not compensate for a process that changes shape every month.
This is the framework's most consequential claim, and it is why scoring matters. An organisation averaging 3.4 across five dimensions and scoring 1 on control is not "moderately ready" — it is not ready, and the average has actively concealed that.
The five dimensions
| Dimension | The question it answers | Level 1 | Level 3 | Level 5 |
|---|---|---|---|---|
| Data | Can the numbers carry the weight? | Sources disagree; reconciliation is manual and undocumented | One system of record per domain; lineage traceable for key figures | Data quality measured continuously; issues detected before they reach a report |
| Process | Is there a stable thing to learn from? | Every preparer has their own way | One documented way; exceptions named and owned | Process measured and continuously adjusted |
| Control | Can we prove what happened? | Controls exist on paper; evidence assembled at audit time | Controls tested on a sample; evidence produced by the system | Continuous monitoring; every transaction tested |
| Capability | Can the team run it? | No one can evaluate a model's output | Named owners who can interpret, challenge and escalate | Finance staff specify and tune models themselves |
| Demand | Is there a real problem to solve? | "The board asked about AI" | A named process with a quantified cost of the current state | A prioritised portfolio tied to measured outcomes |
Data — the one that decides most outcomes
If reported figures cannot be traced to a source, a model built on them will produce confident output that nobody can defend. Data readiness is the dimension most often scored optimistically and most often responsible for the outcome, because finance teams have usually built enough manual compensation that the underlying fragility is invisible from the top of the organisation.
The diagnostic question is not "is our data good?" — everyone says yes — but: can you trace a single figure in the management accounts back through every transformation to its source system, without asking a person? If that requires tribal knowledge, you are at level 1 or 2 regardless of how the data feels.
Process — the one people think they have
Covered in depth in our sequencing work: standardization is one way, defined inputs, named exceptions, and stability when the most experienced person is absent. Most finance functions overestimate this dimension by one to two levels, because familiarity feels like standardization.
Control — the one that arrives late and expensively
Control readiness is rarely the blocker to starting and frequently the blocker to scaling. A pilot needs no control framework. Moving a model into the close does, and the requirement is not merely that the model works but that you can evidence what it did, why, and that a human reviewed it where materiality demanded.
This dimension is where AI differs most sharply from conventional automation. A rules engine does the same thing every time, so testing it once establishes what it does. A model does not offer that guarantee, which means the control has to be over the output and the review, not just the configuration.
Capability — the one that cannot be bought
The dimension most reliably ignored in business cases. Someone must be able to look at a model's output and know whether it is wrong. Not whether it is plausible — whether it is wrong. That is a finance skill, not a technical one, and it cannot be outsourced to the vendor whose product produced the output.
The practical test: when the model proposes an unusual coding, who decides whether to accept it, and on what basis? If the answer is "the system is usually right", capability is at level 1 and the control dimension is compromised regardless of its own score.
Demand — the one that should be checked first
The cheapest failure to avoid. A great deal of finance AI activity begins with a board question rather than a business problem, and produces pilots that succeed technically and are never adopted, because nobody needed the thing they proved.
Level 1 here should stop the programme, not slow it. If no named process has a quantified current-state cost, the correct next step is not a pilot — it is finding the problem.
What to do at each level
A score without a consequence is an assessment, not a framework.
| Weakest dimension at | What this means | The right next move |
|---|---|---|
| Level 1 | Not ready. A pilot will succeed and teach you nothing transferable. | Fix the dimension. This is process, data or organisational work — no product performs it for you. |
| Level 2 | Ready to learn, not to deploy. | Run a contained pilot on a non-critical process, explicitly to test the weak dimension rather than the technology. |
| Level 3 | Ready to deploy narrowly. | Deploy on one process, with human review retained. Measure. Do not scale on enthusiasm. |
| Level 4 | Ready to scale. | Extend across processes, reduce review to risk-based sampling, formalise monitoring. |
| Level 5 | Ready to operate. | Portfolio management: retire what does not earn its place. |
The most common real position is a spread — 3 or 4 on data and process, 1 or 2 on control and capability. That pattern produces the characteristic finance AI failure: successful pilots that cannot be moved into production, because production is where control and capability are load-bearing and the pilot never tested them.
How to run the assessment
Ninety minutes, the right people, and a rule against optimism.
- Score independently, then compare. The controller, the systems owner and the process
owner should score separately. Where they disagree by two levels or more, that gap is the most useful output of the session.
- Score a specific process, not the function. "Finance" has no readiness score. The
close, AP and reconciliation each have one, and they will differ.
- Require evidence for any score above 2. "We think our data is good" is a 2. A traced
figure is a 3.
- Record the weakest dimension and the date. It is the only number that governs the
decision, and the point of the exercise is to move it.
**The strongest objection is that this framework describes readiness for deploying AI into core finance processes, and then applies its caution to everything.**
That is a real overreach in the framing, and it deserves a direct answer. A large share of useful AI in finance today sits nowhere near the ledger: drafting commentary, summarising contracts, interrogating a policy document, accelerating research and first-draft analysis. None of that requires data lineage, control frameworks or model governance, because none of it produces a number anyone books. Applying five-dimension readiness to a controller using an assistant to draft a variance narrative is not rigour — it is obstruction, and it hands the organisation's most capable people a reason to route around finance.
The distinction that holds is whether the output enters the record. If a model's output is reviewed, rewritten and used as an input to human judgement, the readiness bar is genuinely low and the right posture is to enable it quickly with sensible guardrails. If the output becomes a journal entry, a reconciliation certification, a coding decision or a disclosed figure, all five dimensions are load-bearing.
Score readiness for the second category. For the first, the binding constraint is usually policy clarity and tooling access, not readiness — and treating it as a readiness problem delays value for no gain.
Answering the board
When the honest answer is "not yet", the unhelpful version is to say so. The useful version has three parts, and it converts a defensive answer into a plan:
- Where we are, by dimension, with the weakest named and the evidence for the score.
- What the weak dimension is costing us already — data and process weakness are not
AI problems, they are current problems that AI makes visible and expensive.
- What we are doing about it, and when we re-score.
That answer is stronger than an unfounded yes, because it demonstrates the thing the board is actually testing: whether the finance function understands its own operating condition.
- Readiness is five things, and the weakest governs. Averaging the five is the most
common way organisations conceal the dimension that will stop them.
- Control and capability are the usual blockers, and both surface late — after a
successful pilot, at exactly the point the programme has momentum and a budget.
- Score a process, not the function, require evidence above level 2, and re-score on a
date. A readiness score with no re-score is a slide, not a management tool.
Where this leaves you. Score one process across the five dimensions this month. If the weakest is data or process, the next move is not an AI project — it is the work that project would have depended on.