Finance Transformation IntelligenceGovernance, Risk & Controls › Perspective

Agentic Finance — What Changes When Software Acts

>

IndependentEvidence-backedReviewed Jul 2026Sources 54Methodology

Agentic FinanceAI Governance, Risk & ControlDecision Making Under Uncertainty

What this publication is for

>

Questions this answers

  1. What actually changes when software acts rather than suggests?
  2. Why does the existing control model not simply extend to agents?
  3. What are the three questions to answer before delegating a finance action?
  4. What should never be delegated, regardless of how good the technology gets?
  5. How do I let the organisation move without acquiring a control finding?
How to read this.  Layer 1 — Executive summary (4 minutes): the answer and the decision guidance.  Layer 2 — Evidence & analysis: the reasoning, market structure, methodology and references — for controllers, transformation leads and analysts.

Most conversations about AI in finance are organised around capability: what the model can do, how accurate it is, which processes it can handle. That framing misses the distinction that actually governs risk, cost and control.

The line is not capability. It is agency.

A model that proposes an invoice coding and waits is a different kind of system from one that codes the invoice and posts it — even if it is the same model, at the same accuracy. The first is a productivity tool operating inside an existing control. The second is a participant in the control environment, and the control model has to change to accommodate it.

The market's vocabulary for this is still forming, which makes it a good moment to be precise. Vocabulary that settles carelessly is expensive to unsettle.

What actually changes

DimensionAI that suggestsAI that acts
Control pointThe human decision to acceptThe boundary: what the agent may do, and when it must stop
EvidenceThe reviewer's approvalThe agent's decision log, plus proof the boundary held
Error surfacesAt review, before consequenceAfter consequence, in the record
Failure shapeIndividual and visibleSystematic and repeated at speed
Segregation of dutiesIntact — a human still approvesMust be reconstructed in the design
ReversibilityNothing has happened yetDepends entirely on what the action was
What the auditor testsThat review occurredThat the boundary was correctly specified and held

The most consequential row is the first. In a suggesting system, the control is the human decision, and finance already knows how to control human decisions — that is what approval hierarchies, materiality thresholds and segregation of duties are for.

In an acting system, the control moves upstream to the boundary specification. What may this agent do? Under what conditions must it stop and escalate? What is the ceiling on value, volume, or novelty? Those are design decisions made once, in advance, and they govern everything the agent subsequently does. Finance functions have very little practice at specifying controls in that form, because until recently nothing in finance required it.

Why the existing model does not simply extend

Three reasons, in increasing order of difficulty.

Speed changes the error economics. A human coding error affects one invoice and is usually caught downstream. An agent applying a wrong rule affects everything matching that rule, immediately, and the aggregate looks entirely normal because it is internally consistent. Consistency is what makes systematic errors hard to see.

Segregation of duties has to be reconstructed. The principle — no single actor controls a transaction end to end — is foundational and it is not automatically satisfied by inserting a human somewhere. If an agent proposes, executes and reconciles, a human who reviews only the exception queue is not providing segregation; they are providing the appearance of it. Designing genuine segregation around an agent is real work, and it is usually discovered late.

Explainability is not the same as a reason. Most model explainability tells you what influenced an output. An auditor asks something different: on what basis was this transaction treated this way, and was that basis correct? A feature-attribution answer does not satisfy that question. Agents that act inside the record need decision logs that record the rule applied and the inputs it applied to — a design requirement, not a model property.

The three questions before delegating anything

Answer these in order. Any unanswered question means the action is not ready to be delegated, regardless of how good the technology is.

1 · What is the boundary, and what happens at it?

Not "what can it do" but "what may it do, and what must it stop for." A useful boundary is specific: value thresholds, transaction types, novelty conditions, confidence floors. And the stop behaviour must be defined — escalate to whom, with what context, within what time.

A boundary with no defined stop behaviour is not a boundary. It is a hope.

2 · What is the evidence, and would it satisfy an auditor?

Before deployment, not after. The test is concrete: take one transaction the agent will handle, and write out what the audit file would contain. If that file does not include what the agent did, on what basis, what boundary applied, and who reviewed what — the evidence design is incomplete, and the time to discover that is now.

3 · What is the blast radius, and how do we reverse it?

If the agent is systematically wrong for a month before anyone notices, what is the exposure and what is the remediation? This question sizes the boundary. Actions that are cheap to reverse — proposing, drafting, flagging, routing — justify wide boundaries. Actions that are expensive or impossible to reverse — paying, posting, certifying, disclosing — justify narrow ones and hard stops.

Reversibility, not accuracy, should be the primary variable in how much agency is granted. This is the practical heart of the argument.

What should not be delegated

Some finance actions should retain a human decision regardless of capability. Not because software cannot perform them, but because the decision carries accountability that cannot be transferred.

  • Certification. Someone attests that the accounts are correct. That is a personal

accountability, and delegating it to a system misunderstands what certification is.

  • Materiality judgements. What is material depends on context, audience and consequence.

A threshold is a proxy for that judgement, not a replacement.

  • Novel treatment decisions. How to account for something the organisation has not

encountered before. An agent will confidently apply the nearest precedent, and the nearest precedent is exactly what is in question.

  • Anything where the control is the accountability. Approval limits exist so that a named

person is answerable. Automating the approval preserves the workflow and removes the point.

This list is short deliberately, and it is not a list of things AI is bad at. It is a list of places where the human is the control, and removing them removes the control while leaving the process intact — which is the most dangerous version of automation, because nothing appears to have broken.

The strongest case against this

The strongest case against this framing is that it will be overtaken, and that caution expressed in a governance vocabulary tends to calcify into permanent obstruction.

The substance of it: every control argument above assumes an agent whose reliability must be bounded because it cannot be fully trusted. If reliability on a bounded task exceeds human performance — which it plausibly already does for high-volume matching and coding — then insisting on human review is not control, it is ceremony. Worse, it is negative control: a reviewer approving thousands of agent decisions is rubber-stamping, and rubber-stamping degrades the reviewer's attention on the cases that genuinely need it. The control theatre actively reduces safety.

There is a second, sharper version. Finance's existing control model was designed for human error — occasional, random, individually visible. It was never well suited to systematic error either; that is why controls testing samples. An organisation that ported its human-era controls onto agents may simply be encoding an outdated model rather than a sound one.

What survives: the argument for boundaries and evidence does not rest on distrust of the model. It rests on accountability being non-transferable — someone must answer for the accounts — and on reversibility being a genuine, non-negotiable property of the action rather than a judgement about the technology. Those two hold at any capability level.

What does not survive, and should be conceded: the review requirement. Human review of high-volume agent decisions is a transitional control, not a permanent one, and organisations should plan to replace it with monitoring and sampling rather than defend it indefinitely. A framework that treats review as permanent will be wrong within a few years, and will have taught the organisation the wrong habit in the meantime.

Exhibit 1 — Suggest versus act, across seven control dimensions

The practical posture

Neither prohibition nor enthusiasm. Three moves.

  • Classify by reversibility, not by process. Grant wide agency over reversible actions

immediately — drafting, flagging, routing, proposing. Grant narrow, bounded agency over irreversible ones. This lets the organisation move fast where the downside is bounded.

  • Design the evidence before deployment. Write the audit file for one transaction. It

costs an afternoon and prevents the most expensive class of finding.

  • Treat review as transitional, and say so. Set the condition under which review reduces

to sampling — a measured error rate over a stated volume — and hold to it. That commitment is what keeps the control from becoming ceremony.

If you remember only three things
  1. Agency, not capability, is the line that matters. The same model that suggests and

the same model that acts require different control models, because the control point moves from the decision to the boundary.

  1. Classify by reversibility. Cheap-to-reverse actions justify wide agency now;

expensive-to-reverse actions justify hard stops regardless of accuracy.

  1. Accountability does not transfer. Certification, materiality and novel treatment stay

human — not because software cannot do them, but because the human is the control.


Where this leaves you. Take one process where AI is already proposing something, and write the audit file for a single transaction as if the agent had acted instead. That document is the fastest way to find out whether you are ready to cross the line.


Layer 2

Evidence & connections

The reasoning behind the summary above — market structure, methodology, trade-offs and references, for finance transformation leaders, controllers and analysts.

What this rests on

Methodology →
  • The dilynx capability model — journal entry controls, SOX controls and audit, transaction matching, AI invoice coding
  • The dilynx capability maturity spine, in particular the boundary between L3 (Automated) and L5 (Intelligent)
  • dilynx vendor evidence base — AI capability positioning across 71 tracked vendors, with 50 carrying curated AI analysis from cited sources
  • Reasoning and judgement — this is an argued position on a fast-moving question and is presented as such, not as a measured finding
Where a statement is judgement rather than a measured finding, it is labelled as such in the text. Independent — no paid placements. Rankings are never influenced by commercial relationships. Our independence →

Related benchmarks

Benchmark Intelligence →

How performance in this area is measured, and what comparable finance organisations achieve.

Close & ReportingAutomation & AI Adoption

Related implementation

Transformation Marketplace →

Where the answer is a partner rather than a product — the specialisms that deliver work in this area.

AI Strategy for FinanceFinance AutomationFinance Operating Model

Related assessment

How it works →

The Executive Finance Assessment reads your organisation against the same maturity spine, decision archetypes and benchmark models used across this pillar — so what you read here and what it tells you about Agentic Finance — What Changes When Software Acts are expressed in one vocabulary, not two.

Executive Finance Assessment

What does this mean for your organisation?

This research frames the question in general terms. The Executive Finance Assessment answers it for your finance function specifically — your position, your highest-impact move, and the evidence behind it.

Begins with a free Executive Brief — about five minutes, anonymous, no account. Full assessment €59, one-time. It complements the research; it does not replace it.