The Signature Problem: Why SMCR Is the Best Thing That Will Happen to Your Agentic AI Programme
Don't let human-in-the-loop turn into late-night whack-a-mole
You are already personally accountable for how agents are governed in your domain. The FCA's Mills Review and the Treasury Committee confirm it: you cannot wait for a new rulebook, because the existing Conduct Rules apply to agents today.That accountability is not a burden to delegate to your technology teams — it is your mandate to lead them. In the SMCR frame, the regulator has named you, not your engineers. The operations leaders who act on that will own how agents enter their processes: triggers that force evaluation rather than wait for suspicion, judgment concentrated on exceptions rather than diluted across the routine, and a continuous evidence pack that answers the regulator before they ask — even when the model you tested changes beneath you without notice.The firms whose boards say yes to the next deployment will be the ones whose operations leaders can prove control of the last one. Speed follows defensibility.
You have already signed
In July 2026, the FCA published the Mills Review, the most substantial regulatory statement yet on AI in retail financial services. One sentence in it should be pinned above the desk of every senior manager in operations: "It will not be enough to say that a person remains 'in the loop.' Firms will need to be clear about what the person is expected to do, what information they receive, when they can intervene, how challenge is recorded and how escalation works."
Notice what that sentence assumes. Not whether your firm will deploy AI agents in regulated processes. Not whether accountability should attach to them. It assumes both, and asks only what your defence looks like. The Review is explicit that the existing framework — the Consumer Duty, the Senior Managers Regime, operational resilience — remains the basis of regulation. You cannot wait for a new AI rulebook; the existing Conduct Rules apply to agents now. Which means that if an agent is chasing payments, triaging complaints, or clearing exceptions anywhere in your domain, you are personally accountable for the reasonable steps that govern it today, under a signature you have already given.
The Treasury Committee heard the same point in January, put more bluntly. David Geale, the FCA's Executive Director for Payments and Digital Finance, told MPs that the existing regime already ensures individuals are "on the hook" for harm caused to consumers through AI — and he rebutted the industry argument that AI's "lack of explainability" makes personal accountability unworkable: "'I did not understand it' is not a defence, because you should understand what you are deploying, or you should understand what you are seeking to achieve." No new senior manager function is coming, he added, because the existing framework already captures it.
This is the part most commentary gets wrong. The dominant industry response — a wave of client alerts from law firms and consultancies — reads this as threat. Fines. Prohibition. Public censure. The message is: be afraid, then buy something.
We think that framing is not only lazy, it is strategically backwards. The senior managers who read SMCR as a threat will spend the next three years defending pilots they cannot explain. The ones who read it as a specification will spend them deploying agents their boards actually trust.
The failure nobody will see coming
Start with the uncomfortable finding at the centre of our own research: override without a trigger is ineffective oversight. This is not a line from any regulator. It is the conclusion of Null Proof's internal investigation into SMCR accountability for agentic AI, and the Mills Review's analysis points the same way: oversight weakens when reviewers are overloaded, lack the right information, or rely too readily on model outputs. Where the Review stops at the diagnosis, our work has been on the mechanism — and the mechanism is where personal accountability is won or lost.
Most agent deployments today come with a reassuring slide — "human in the loop." A named person can, in principle, review and override the agent. The problem is that "in principle" collapses under four documented failure modes:
- Automation bias. When the agent's output looks correct — and at scale it almost always looks correct — humans do not question it. The intuition to intervene is suppressed precisely when the system is working smoothly.
- Scale mismatch. An agent processing thousands of decisions a day cannot be supervised by a person who can meaningfully review dozens. Human-initiated review is arithmetically impossible at agent-scale volume.
- Drift blindness. Agents degrade gradually. A model whose behaviour shifts by small increments never crosses the threshold of any individual's attention — until the accumulated drift is the incident.
- Novel failure modes. Human intuition is trained on human-scale failures. Agentic systems fail in ways nobody has seen before, and intuition, by definition, cannot anticipate what it has never encountered.
The regulatory consequence is sharp, and it lands in two places at once. Where the EU AI Act applies — for firms operating in or serving the EU market — Article 14 requires an overseer to demonstrate four capabilities: understanding the system's capabilities and limits, awareness of automation bias, correct interpretation of its output, and the ability to override or halt it. Map those against SMCR Conduct Rule 2 — the obligation to take reasonable steps to ensure your business complies — and the two regimes converge on the same test, from opposite directions. The AI Act's obligations and penalties attach to the undertaking: ineffective oversight can put the company's balance sheet in play, with administrative fines for breaches of the oversight obligations reaching €15M or 3% of global turnover. Your exposure runs through SMCR, where the oversight failure may be evidence that you failed to take reasonable steps. Corporate liability on one side, personal accountability on the other — and the same artefact, or its absence, decides both. Note, too, that the Act requires deployers to assign oversight to people with the necessary competence, training, and authority (Art. 26(2)): the regulator expects a named, capable human by design, not by org-chart accident. The "human in the loop" slide is not a defence. Without a mechanism that forces the loop to close, that slide becomes evidence that oversight existed in name only.
The inversion: regulation as the design spec for autonomy
This is the inversion. Read the obligations again, not as constraints but as requirements engineering.
The regulators are not asking for a human watching a screen. They are asking for an architecture of forced evaluation — a system in which the senior manager's attention is compelled, on evidence, at the right frequency, by the system itself. That is not a compliance burden. It is the first credible specification we have seen for what trustworthy agentic autonomy requires. Any firm serious about running agents at scale needs exactly this architecture whether or not a regulator ever asks for it. SMCR simply makes the need personal, dated, and evidential.
There is a harder implication in this, and it changes what being accountable means in practice. If the oversight architecture is what your signature defends, then that architecture cannot be something you receive. It has to be something you specified. An agent built elsewhere and shipped into your domain arrives with its triggers, thresholds, and escalation paths already decided by people who do not carry your accountability. Reviewing it at UAT — reading a report a week before go-live and signing a box — is not oversight of the system. It is oversight of a decision someone else made about what you would be allowed to see.
The accountable person therefore has both an obligation and an opportunity to be inside the design, driving the requirements: which anomalies force evaluation, what confidence scores must mean before a decision proceeds, where the circuit breakers sit and what trips them, which exceptions escalate and to whom, what the evidence pack captures from day one. None of these are technical details to be delegated to the delivery team. They are the substance of your reasonable steps. A senior manager who shaped them can answer the only question SMCR asks — what reasonable steps did you take? — with evidence of their own agency. A senior manager who received them can only describe other people's steps, and under this regime that is not evidence of control.
In practice, the architecture has four layers, each forcing a different kind of evaluation:
| Layer | Mechanism | Frequency | What it forces you to see | |---|---|---|---| | Real-time | Anomaly detection, threshold breaches, circuit breakers | Continuous | The event you would never have noticed | | Scheduled | Calendar-driven performance and trend review | Weekly / monthly | Drift that no single event reveals | | Adversarial | Red-team and edge-case testing | Quarterly | Failure modes nobody has encountered yet | | Strategic | Capability validation against declared purpose | Annual | Whether the agent still does what you signed for |
Note the property these layers share: none of them depends on the senior manager deciding to look. That is the entire point. Intuition is the thing that fails; the architecture is what replaces it.
The operating model: algorithm, exception, review
The architecture answers the oversight question. A second model answers the operational one — and it resolves a tension most operations leaders feel but rarely articulate: how do you let agents run without abandoning judgment?
The answer, which maps cleanly onto Conduct Rules 1 through 3, is to remove human judgment from the happy path entirely, and concentrate all of it on exceptions.
Consider late-payment chasing — a process drowning in exactly the subjective, inconsistent, bias-prone human decisions regulators dislike. The defensible model has three components:
- The happy path is deterministic in its decisions, agentic in its execution. This is the point most deployments miss, in both directions. Routine chasing runs on explicit rules — data-driven, consistent, auditable — that a person can read, test, and sign. But this is not the old linear workflow: agents orchestrate those rules dynamically, running parallel tasks, looping, and self-correcting through inputs that match the expected shape in principle if not in exact syntax — the messy variance that breaks rigid automation. The agent never decides: the moment judgment is required, it routes. Rules can be discriminatory too — no one who has lived through a pricing review doubts that — but they are deterministic, and that changes the evidence question entirely. You can test a ruleset once and rely on that test remaining valid when you need to present it: the artefact you validated is the artefact that is running. The same is not true of an API-provided language model, whose behaviour can shift beneath evidence you have already gathered — every update quietly retiring your prior assurance. Judgment out of the loop therefore removes not the possibility of bias, but its opacity and its drift: what remains is inspectable, testable, and stable under scrutiny. You do not audit a rule the way you audit a model. You review the rule — and a quarterly review of a ruleset is something a senior manager can actually do, and evidence.
- All the judgment concentrates in exceptions — made by humans, prepared by agents. The debtor disputes the debt. Vulnerability indicators appear. A payment plan needs negotiating. A legal threat arrives. These are the judgment cases, and the call belongs to a person — but not a person hunting through six systems under time pressure. The agent's job at the boundary is assembly: pulling the debtor's history, lawfully held vulnerability flags, complaint record and prior correspondence into the context the human needs to judge well. The routing rule that gets a case there is explicit, tested and reviewable — not someone's inbox habits — and whatever model risk remains lives at that boundary: small, specific, and testable.
- Quarterly review closes the loop — but it is not the only loop. The triggers from the previous section govern the runtime: continuous, automatic, unmissable. The quarterly review governs the rules: is the ruleset still achieving target recovery rates, is it still compliant, are exception volumes and patterns signalling that the boundary sits in the wrong place, and is the human exception handling of good quality. A quarterly review of behaviour at agent scale would be theatre; a quarterly review of an explicit ruleset, backed by continuous triggers, is governance. Rising exception rates are not a nuisance; they are the signal that the rules no longer fit reality — and the review is where that signal lands on a named desk.
Agents absorb variance in execution; humans absorb variance in judgment. The result is a division of labour in which humans do less work but carry more judgment — and, critically, in which every element the senior manager is accountable for has a corresponding artefact. That correspondence is the whole game, because SMCR does not ultimately ask whether you were careful. It asks whether you can show you were careful.
The evidence pack is the defence
This is the part that separates firms that talk about governance from firms that can prove it. Under SMCR, the "reasonable steps" defence is evidential or it is nothing. No artefact, no defence — regardless of how diligent you genuinely were.
The evidential spine of a defensible deployment is mundane, continuous, and unglamorous:
| Evidence | Frequency | What it proves | |---|---|---| | Override decision log | Continuous | You can and do halt the system (Art. 14(d)) | | Circuit breaker activation log | Per event | Triggers fire and you respond | | Performance review sign-off | Weekly / monthly | Scheduled evaluation happens | | Audit trail completeness report | Monthly | Decisions are reconstructable | | Red-team and override test results | Quarterly | You probe for failure, not wait for it | | Capability and competence validation | Annual | You understand what you signed for |
If a regulator asked tomorrow for the evidence that your oversight of the collections agent is effective, could you produce it at once? For most firms running pilots today, the honest answer is no — not because they are reckless, but because nobody built the pack. The gap is architectural, not moral. Which also means it is fixable, fast, by design rather than by remediation.
One more uncomfortable line item belongs here. Any FCA penalty imposed on you is yours by design: firms are prohibited from paying FCA penalties imposed on their senior managers, and such fines are largely uninsurable as a matter of public policy. What insurance is for is the defence — the legal costs of an FCA investigation and enforcement process, which escalate quickly into material personal expense in contested cases. So review your D&O cover for AI exclusions with your broker or adviser before your next renewal. If the policy does not respond, it may be your defence costs in that investigation — not the company's P&L — that are left exposed.
The licence to go fast
And so to the conclusion that the fear-mongering version of this story never reaches.
The question was never "can we deploy agents under SMCR?" The FCA answered that by declining to write new rules: the framework you already operate under was the answer. The real question is older and harder: can you defend the signature you already gave?
The senior manager's signature has never been a brake on the business. It is the mechanism by which a regulated firm is permitted to operate at all — the licence, renewed daily, to take risk at scale. Agentic AI does not change that logic; it intensifies it. The firms that build the trigger architecture and the evidence pack first will not merely be compliant. They will be the only firms whose boards, confronted with the next proposed deployment, hear "yes, and here is the proof of how we control the last one" instead of "yes, and we believe the vendor's slide about humans in the loop."
Speed follows defensibility. In a regime of personal accountability, that is not a paradox. It is the entire opportunity.
Sources
- FCA, The Mills Review: AI and the future of retail financial services (July 2026) — Source of the "in the loop" passage on human oversight (System Shift 1, p.30) and the observation that oversight weakens when reviewers are overloaded, under-informed or over-reliant on model outputs (p.17). Confirms the existing framework (Consumer Duty, SMR, operational resilience) remains the basis of regulation, with SMR accountability an emerging pressure point at Approver/Observer autonomy levels. Landing page: fca.org.uk.
- House of Commons Treasury Committee, AI in Financial Services (January 2026) — Source of David Geale's evidence (para 12): individuals "on the hook" for AI-caused consumer harm; "'I did not understand it' is not a defence, because you should understand what you are deploying"; no new senior manager function required. Quote verified manually against the report; automated fetch blocked (403).
- EU AI Act, Regulation (EU) 2024/1689, Articles 9–15, 26, 43, 99 — Human oversight (Art. 14), administrative fines (Art. 99), deployer obligation to assign oversight to persons with necessary competence, training and authority (Art. 26(2)), risk management, logging, and conformity assessment obligations mapped against SMCR reasonable steps.
- Aveni, SMCR, AI agents and senior manager accountability (June 2026) — "SMCR compliance has always been about one thing: making sure a named individual can demonstrate they took reasonable steps."
- Bank of England, Summary of AI roundtables (February 2026) — "Conventional approach to model risk management… is becoming unworkable"; supports shift from periodic to continuous oversight.
- SSRN (2026), on the goal–plan–execution gap in agentic systems — Humans do not spontaneously detect when agents go off track; basis for systemic trigger argument.
- ASI Online, The Authority Gap: Why Human Oversight Fails the Agentic Workforce (2026) — Human-in-the-loop triggers automation bias; passive oversight is ineffective oversight.
- Springer, AI & SOCIETY (April 2026), on agency and alignment — Human evaluative agency must be structured, not reactive.
- Black-box auditing of production LLM endpoints, arXiv 2605.29524 (2026) — Longitudinal probes of 16 production endpoints over 64 days; one endpoint crossed the drift threshold at week 9, shifting one-directionally in a manner consistent with a silent backend update rather than sampling noise. Basis for the silent update callout.
- Anthropic, Claude Desktop changelog — Release cadence observed 28 July 2026: 16 builds in 30 days, including v1.24012.0 (21 July) and v1.24012.9 (24 July).
- OpenAI, Codex CLI releases (GitHub) — Release cadence observed 28 July 2026: 63 releases in 30 days (11 stable, 52 prereleases).
- ToSea.ai, Verifying LLM API models by fingerprint (2026) — Practitioner write-up citing CISPA's endpoint verification work ("Real Money, Fake Models") and the aggregator audit finding that 11 of 31 commercial endpoints serving Llama models produced output statistically incompatible with vendor reference weights. Basis for the aggregator substitution claim in the silent update callout.
Underlying research: Null Proof internal investigation, "FCA SMR × EU AI Act × Agentic AI — Individual Accountability Investigation" (July 2026), with verified source table. Sources 6–7: URLs verified; content fetch blocked (403) at time of writing.
This paper is signed with Ed25519. Run the signature check in your browser to confirm it hasn't been altered since it was signed.