The question agentic AI forces
Organisations introducing agentic AI soon meet a fundamental question. When AI can analyse evidence, recommend actions and increasingly execute work, how does the organisation determine which decisions machines are authorised to make and which require accountable human judgement?
That is an architectural question, not simply a matter of policy. A policy can say a person must approve consequential decisions. The architecture has to know which decisions those are and who may make them.
The answer matters well beyond governance. Keep every existing approval and handoff, and AI may speed up tasks while the operating model stays as it was. Remove them without explicit decision rights, and machines may exercise authority nobody intended to delegate. Making authority explicit and executable escapes that trade-off. It lets a bank redesign work around the decisions that genuinely require accountability, rather than the processes that historically implemented it.
My perspective comes from more than 25 years of delivering complex banking technology and transformation programmes, and from building Henova, an AI-native health company. There, AI participates in consequential work rather than simply assisting people. Healthcare is not banking, but both have consequential decisions, established accountability, regulatory obligations, risk classification, specialist authority and demanding requirements for evidence and audit.
The point is not to copy clinical rules into banking, but to share the architectural lessons that transfer. Here, the lesson is this: the more capable AI becomes at reasoning and executing work, the more explicitly the organisation must separate intelligence from decision-making authority.
Consequence shapes the workflow
A concrete illustration comes from how Henova governs changes to its own software. Before work is built, it is given a clinical safety classification describing what could happen to a person if that software failed.
Class A means no injury is possible. Failure cannot reach a health decision or mask anything safety-relevant, as with a layout change or internal tooling with no clinical content.
Class B means non-serious injury is possible, because the user might be misinformed. Displaying a biomarker value is an example. The system presents information; it does not judge the individual.
Class C means failure could contribute to serious injury. It covers software that applies clinical judgement, classifies individual risk, generates recommendations or controls safety-critical outcomes.
The line between B and C is instructive. "Your LDL-C is 145 mg/dL" is a factual display. "Your LDL-C places you in this risk category" is an individual clinical judgement. The number is the same; the consequence is different.
The grade isn't just metadata. It changes the executable workflow and determines where human authority is mandatory. True Class A work can follow a highly automated route if its other controls are satisfied. Class C work meets boundaries that cannot be crossed without an authorised clinical ruling. Because work can change during implementation, the grade is reassessed as it evolves.
This classifies software changes, not every decision the product makes at runtime. Accepting a safety-relevant change and ruling on a consequential operational decision are different workflows, resting on the same principle: consequence determines where accountable authority must be exercised before work proceeds.
The significance is commercial as well as clinical. Different consequences require different levels of authority and intervention. Once that is designed into the workflow, lower-consequence work can be automated without people supervising everything.
A bank faces comparable questions when AI participates in lending, financial advice or risk acceptance, or changes the systems that implement those decisions. The classifications and regulations will differ. The questions will not. Who has authority? Which decisions can be delegated? What evidence is required? What prevents execution before that authority has been exercised?
"Human in the loop" says too little
The phrase is reassuring and architecturally almost empty. It leaves the important questions open:
- Which human?
- Deciding what, and on what evidence?
- At which point, and with what authority?
- What happens next?
"A clinician reviews it" or "a credit officer signs off" hides all of them. Review can mean glancing, commenting, approving or owning the outcome. A system that cannot tell the difference cannot know whether it is permitted to proceed.
Authority is not intelligence
Agents are not task-runners. They reason, analyse, recommend, challenge, assemble evidence, plan and execute. They can do a great deal of the thinking.
But intelligence is not authority. Where accountable judgement is required, the decision belongs to someone with the right to make it.
So human involvement is targeted at authority boundaries, not spread across the work. The aim is bounded autonomy, not manual approval of every AI action. At a boundary, the system should know:
- that a human decision is required;
- which person or body holds the authority;
- the question to be decided and the evidence available;
- the scope of the ruling;
- what may happen once it is made.
Humans are authorities within the system, not an emergency fallback.
The ruling must become state
This is where governance becomes architecture.
A decision that produces "Approved" is not enough. Nor is a record that says "We discussed this with Sarah and she was happy." Neither tells a machine what it may now do.
If a human judgement governs subsequent machine behaviour, the judgement has to become something the machine can reliably consume. That means recording what was decided, who ruled and under what authority, the scope and conditions of the ruling, the evidence it rested on and which earlier decision it supersedes.
The result is a lineage: evidence, authorised ruling, governed state, downstream machine action. That is materially different from an audit log recording that someone clicked a button.
Once a human has ruled, an agent should not reopen the judgement simply because it can reason about the same evidence. Agents may challenge where the architecture permits it and new evidence justifies it. They may not silently take the decision right. Confidence is not authority.
Banks already have the concepts
Banks hold many forms of human authority that will need this treatment.
Data governance is the clearest example. An agent can find conflicting definitions of "customer" or "default" and recommend an authoritative source. The accountable data owner settles the definition. If that ruling lives only in committee minutes, every downstream agent has to rediscover it. As governed state, they all act against the same one.
Risk acceptance and policy interpretation follow the same pattern. Agents assemble evidence and precedent. The accountable risk, Compliance or business owner rules, and machines execute inside that ruling.
Delegated authority, maker-checker, segregation of duties and escalation are already familiar in banking. The new problem is not inventing those concepts. It is making them executable where machines increasingly reason and act.
Redesigning work around authority
The question for a bank is not whether to keep its controls. It is which of them express genuine accountability, and which are historical mechanisms for implementing it.
Preserving every approval and handoff buys task productivity, not operating-model change. Removing them without explicit decision rights lets machines exercise authority the organisation never delegated. Neither is a transformation.
The alternative is to redesign work around explicit authority boundaries. Business leadership decides which decisions can be delegated, within what limits, and which judgements require accountable human authority. Technology makes those boundaries executable, so the right questions reach the right people with the right evidence and their rulings become state that machines act on.
Agents then take on more analysis and execution inside those boundaries, while people concentrate on decisions that genuinely need their authority. Done well, the result can be greater automation, fewer unnecessary handoffs and more deliberate use of human judgement, without abandoning accountability.
These are potential benefits of an operating model designed this way. Henova has made the architectural problem clear at small scale; it has not proven these outcomes for a bank. Realising them starts with leadership deciding where authority sits.
Humans arbitrate. Agents execute. The architecture has to know the difference.