The central message
AI is changing everything, except the regulation
The Review's conclusion is surprisingly simple:
The existing regulatory framework works. There is no need for a new AI rulebook.
Nobody argued that the Senior Managers Regime should change. Firms weren't asking for new rules. They were asking to be told how the current ones apply.
I've already heard this reported as good news, and it is, but it's also a double-edged sword, especially if you're an executive. There is no implementation deadline, no new consultation to respond to, no external event that forces this onto your board agenda, and that all means there is also no reason left to wait.
No deadline does not mean no obligation.
"We are waiting to see what the regulator says" was a defensible position in in H1 2026. It isn't any longer. What this means is that the rules were live the whole time. So if you're already using AI, the urgency is now operational rather than regulatory.
For your firm you need to know:
- what AI is being used (officially and unofficially);
- Where and how it is being used;
- what decisions it influences;
- who owns each use case;
- what controls are operating around it; and
- what happens when the system changes e.g. a software provider enables an agent.
1. Risk is moving up the autonomy spectrum
The bit worth getting your head around
The single most tangible takeaway from the Review is its five-level autonomy spectrum.
What makes it useful is that it defines the levels by what the human is doing, rather than by how clever the machine is.
At Level 1, a person does the work and AI helps. At Level 5, AI does the work and a person watches.

From this, the Review then maps the existing regulatory regimes against those levels and shows where each starts to strain, which of course varies. The perimeter question, whether something needs a license at all, strains earliest. The advice boundary and operational resilience come under pressure in the middle. The Senior Managers Regime and Consumer Duty hold up until the top, where proving that a named human was meaningfully in control becomes genuinely difficult.
That is important because it forces you away from one of the biggest issues we've seen in firms which is doing one AI risk assessment across the whole organisation.
A chatbot that tells customers your opening hours and a system that drafts lending decisions are not remotely the same thing.
From this, the Review then maps the existing regulatory regimes against those levels and shows where each starts to strain, which of course varies. The perimeter question, whether something needs a license at all, strains earliest. The advice boundary and operational resilience come under pressure in the middle. The Senior Managers Regime and Consumer Duty hold up until the top, where proving that a named human was meaningfully in control becomes genuinely difficult.
That is important because it forces you away from one of the biggest issues we've seen in firms which is doing one AI risk assessment across the whole organisation.
A chatbot that tells customers your opening hours and a system that drafts lending decisions are not remotely the same thing.
Different risk. Different regulatory position. Different control requirement.
But there's another shift here that is even more important, and novel to the world where AI is now often embedded in core systems by software providers.

FIGURE 2 — The autonomy spectrum
AI risk is progressive, not static
Your risk isn't simply where a tool sits today, it is where it can move. Your supplier ships an update, the tool gets better and it starts doing more on its own. It has quietly climbed the autonomy spectrum and your assessment is now out of date.
This is a fundamental shift in how suppliers and systems have historically been managed.
The thing you use is no longer necessarily the thing you assessed.
That means every material AI use case needs:
- a named business owner;
- a defined autonomy level;
- clear evidence of what the system is allowed to do; and
- a trigger for when the assessment needs to be revisited.
THE EXECUTIVE TEST
If this system became more autonomous tomorrow, how would we know, and who would be responsible for reassessing it?
2. "A human checks it" is no longer a control
The Review is unusually blunt here. It says that it is not enough to say that a person remains "in the loop."
Firms need to be clear about what that person is expected to do, what information they receive, when they can intervene, how challenge is recorded and how escalation works.
p26, Mills review
That sounds obvious, but in practice, it is a major change especially as human in the loop has been the default solution to the limitation of LLMs across accuracy, reliability and auditability.
The review is effectively saying that the existence of a human is not the control, but what the human actually does is the control.
Oversight collapses when reviewers are overloaded, when the machine sounds confident or when the reviewer doesn't have the information or authority needed to challenge it.
I co-wrote a longer piece on the risks and issues of human-in-the-loop, Get Out of the Loop, for anyone who wants the empirical detail on why human-in-the-loop is statistically more likely to be performance art than effective control. It's been proven across all sectors over decades, we were just hopeful it might solve a gap with LLMs we don't know how to address otherwise.
But back to the Mills Review. Questions to ask here could include:
If you were not allowed to describe something as "human in the loop", how would you describe the control?
What does the reviewer actually see on screen? What information do they receive?
What are they empowered to challenge?
How often do they say no? What happens when they do?
If the answer is that they see a recommendation, a confidence score and a queue of 400 more, they aren't checking anything. They're clicking, and you are paying salaries for a control that exists on paper but doesn't actually work. It's not a compliance failure; it a core design failure.
THE EXECUTIVE TEST
If we say there is human oversight, can we show what that oversight actually looks like in practice? And can we audit how oversight happened and correct it?
3. Validating AI once is over
Almost nobody validates a model more than once. You test it, sign it off, switch it on and, at best, review it annually.
That worked reasonably well when the thing being assessed stayed relatively stable. It is much harder to defend when the system changes after you buy it, and nobody sends you a note when it does, as mentioned above. The so-what highlighted in the Review is that the expectation is moving towards continuous monitoring. To deliver to this you need to be able to say:
- what you monitor;
- what threshold matters;
- who sees the signal;
- what happens next; and
- who has the authority to intervene.
And you cannot sensibly rely on the model itself to do, which is the other scenario we see creeping in. We cannot reply on models to mark their own homework.
Don't ask the model to explain itself
Generative AI will happily produce an explanation of its reasoning and whilst it may sound convincing any explanation you get is not a record of how the system actually reached its conclusion. It is a reconstruction; a plausible story about how it might have got there.
That is not the same thing as independent evidence, and it is not auditable. Anyone who says otherwise has not understood the core mechanics of how LLMs work.
Don't ask the model to watch itself
A model has no reliable sense of when it is wrong, so it can be every bit as fluent when it is hallucinating as when it is right. In other words, it cannot tell if something is a hallucination or is correct. So a system with those blind spots cannot reliably be the control that detects them.
That leaves you needing monitoring that sits outside the model and does not depend on the model's own account of itself. It is also why "we will monitor it continuously" appears in a lot of AI governance policies and in very few operating models.
THE PRACTICAL QUESTION
What signal would tell you that the system has changed or degraded before a customer or regulator does?
4. Independence is the thing you are actually buying
There is a perception doing the rounds that if you run two models and have them check each other, you've built yourself a control. But you haven't necessarily. A control only works if it fails independently of the thing it is controlling; that is the principle behind a second line and two models do not automatically satisfy it.
They are trained on the data that happens to be available rather than all of it. For a specific regulated task, the available pool is narrower still. Two models from two vendors, aimed at the same job, may be reading from near enough the same library and so can inherit the same gaps and biases. So when they fail, they can fail together and for the same reason.
It's the same opinion twice, with a second invoice.
Something to keep in mind when you read the Review's warnings about concentration.
Firms leaning on the same small set of providers may start behaving in correlated ways: making synchronised pricing or underwriting decisions and withdrawing from the same customer segments at the same time. It's the same failure just one layer up.
Shared dependency destroys independence.
Whether the shared thing is a model inside your stack or a provider underneath half the market.
The failure mode I'd actually test for
Most firms have planned for a supplier going down but very few have planned for a supplier that stays up and quietly gets worse. Degradation rarely raises no alarm or produce a red status on a dashboard, it just makes marginally worse decisions for several months while everything appears to be working.
You won't detect that with an availability metric or by asking the system how it thinks it is doing.
THE OPERATION QUESTION ISN'T
What happens if the supplier fails?
It's: How would we know if the supplier is slowly failing?
5. AI changes the operational risk profile
The Review contains several findings that, taken together, point to a broader shift. AI doesn't necessarily create entirely new risks. It changes the speed, scale and economics of existing ones.
Fraud gets cheaper, not cleverer
The Review's assessment is that AI amplifies the fraud we already have rather than inventing a new category of it. Cloned voices, synthetic identities and convincing fake documents are now cheap to produce and simple to run at volume.
The skill barrier has gone. The population capable of running a persuasive scam has therefore expanded considerably, and that changes where the operational strain falls.
When attempts were expensive, detection was the hard part. When the cost per attempt approaches zero, volume becomes the problem and the bottleneck moves downstream:
Detection → triage → investigation → reporting
The Review specifically identifies suspicious activity reporting as a point that can jam. If your financial crime team is already at capacity, that's the number I'd model.
THE COO QUESTION
Can our operation absorb a large increase in attempts without simply moving the bottleneck from detection into investigation?
There is another second-order issue worth registering now.
As customers begin transacting through agents, some fraud will occur between machines with no human anywhere in the exchange. Controls designed to detect a person being deceived have nothing to work with in that scenario. This isn't an immediate problem for most but it does mean your current detection strategy has a shelf life.
Your cyber exposure is probably duller than you think
The Review anchors its cyber analysis on the AI Security Institute's evaluation of a frontier model, which found it capable of exploiting systems with weak security posture.
The AISI's caveat, reproduced in the Review, is that it cannot say whether the model could breach well-defended systems. That caveat is the whole finding.
AI isn't necessarily creating novel vulnerabilities, it is getting better at locating existing ones faster and at scale. Which means your exposure is not exotic.
It is:
- unpatched software;
- incomplete asset inventories;
- slow remediation;
- legacy estates; and
- complex outsourcing chains.
THE PRACTICAL QUESTION
How long does it take us, in practice, from a vulnerability being disclosed to it being fixed across our estate?
If the answer is measured in months, the Review's implication is uncomfortable. The gap between adequately and inadequately defended firms is about to widen sharply.
There is also an upstream problem. An incident at a widely used AI provider lands on every firm relying on it at once and that makes supplier concentration a resilience question, not just a procurement question.

6. Your data estate is the ceiling
The Review cites Cambridge's finding that data quality is the single largest obstacle to scaling AI in financial services.
For regulated decisions, the standard is clear: accurate, complete, timely, consistent and traceable data, with lineage from input through to outcome.
A useful answer is not enough if the basis for it cannot be reconstructed
p17, Mills Review
Most firms cannot meet that today, and there is a good reason. Data estates were built to produce periodic returns. They weren't built to support constant machine decision-making that can also be reconstructed afterwards.
Those are different engineering problems. Reporting can tolerate reconciliation after the fact but evidencing a decision cannot.
This constrains you twice. First, it caps what you can responsibly deploy; a decision you cannot trace is a decision you cannot properly defend. Second, it weakens the regulator's confidence in what you submit which now matters more than it used to.
The work is slow, expensive and almost completely invisible to customers, which is precisely why it keeps losing to more legible projects, but is nonetheless the gating item.
A useful diagnostic
Take one decision and one customer. Trace it end to end from source data to outcome.
Then ask:
- Who owned each step?
- What changed?
- What was recorded?
- What evidence remains?
- Could you reproduce it six months later?
How cleanly you can do that tells you more about your readiness than almost any maturity assessment.
7. The regulator is changing shape too
Recommendation 6 proposes an AI-enabled supervisory model.
The FCA would take its own systems to the point where they recommend and prepare, while human supervisors retain the decisions, explicitly stopping short of full autonomy.
The stated capabilities include:
- monitoring outcomes across firms in near real time;
- aggregating complaints and pricing data across the market; and
- spotting patterns invisible from inside any single firm.
The Review gives three strategic reasons for this shift.
1. Greater efficiency
AI can accelerate authorisation, supervision and enforcement through faster application processing, real-time monitoring and automation.
2. More effective action against harm
AI can identify risks earlier, enable faster intervention and allow the FCA to handle more cases.
3. Keeping pace with technological change
AI will become increasingly central to how regulated firms operate. The FCA therefore needs comparable capabilities if it is going to supervise firms effectively. Taken together, this creates two important implications for firms.
Comparison becomes continuous
If you are an outlier on pricing or complaints, that may surface before you've identified it yourself. You lose the opportunity to frame the issue in your own words. The periodic return has always functioned partly as a narrative device. That utility is diminishing.
Reporting becomes an operating-model question
The Review is candid that this model requires firms to supply structured, timely, high-quality data. It implies a move away from periodic returns towards continuous, event-driven flows. Larger firms will be expected to get there first. That means the next time you invest in reporting infrastructure, the specification shouldn't simply be accurate and on time.
It should also be machine-readable and continuously available.
THE OPERATIONAL QUESTION
Can you see the same patterns internally before the regulator does?
8. The wild west of AI agents
Every control you own rests on knowing who you're dealing with.
Onboarding. Mandates. Authentication. Fraud monitoring.
All of it assumes identity is established at the front door. Twenty years and a great deal of money have gone into that assumption.
AI agents break it.
Not through malice, there is simply no established way to verify one. The UK does not currently define what an agent's identity is, require anyone to verify it or maintain a register of who operates them.
So when something arrives asking to open an account or move money on behalf of Mrs Jones, you have no reliable way of establishing whether it has any business doing so.
Consent doesn't solve the gap. Open banking assumes a person clicks approve while the thing is happening. An agent may act on that instruction a fortnight later, at four firms simultaneously, at three in the morning.
Permission granted once. Exercised repeatedly. At times and places your customer never contemplated. Nor is this necessarily a 2030 problem.
A fifth of UK adults already say they would let AI act for them within limits they set, rising to 28% among those already using it.
The Review therefore suggests that the FCA should start defining:
- accountability and liability frameworks;
- standards for evidencing control and auditability;
- minimum conditions for those acting on behalf of consumers;
- valid instructions;
- structured and verifiable mandates, including scope, limits and revocation; and
- requirements for identity linkage.
For firms using, especially customer-facing, agents, or planning to, deliberate decisions will be critical. Accept, block or accept on conditions.
9. The bigger picture
The future (of AI) is bright
I was heartened by how optimistic the Mills Review is about AI and the positive impact that, if understood and deployed properly, it can have in financial services, particularly for consumers.

I was equally pleased that the Review is explicit about the fact that AI is evolving. We are still at an early iteration of AI in financial services. Today's focus is heavily on LLMs, but LLMs are only one subset of the AI family.
The review is clear on the limitations of LLMs in financial services and that some AI deployed today may not be fit for purpose.
But the rapid development of solutions across AI more broadly means there are technologies in flight that could move the dial, both more safely and reliably.
The whole of section3: The systemic driver: Advances in AI capability is dedicated to this. I will cover this in more detail in my follow up article.

The one key question
What surprised me most personally, though, was how many of the Review's findings resolve to the same underlying question.
Can you show what happened, on whose authority, and why?
"A useful answer is not enough if the basis for it cannot be reconstructed."
That, to me, is the real message of the Mills Review. I have said before that you don't need an AI strategy any more than you needed an electricity strategy or an internet strategy. You need a business strategy and the right tools to deliver it. The Mills Review doesn't change that, but what it does is tell you what those tools need to be able to do.
They need to reliably evidence what happened, on whose authority, and why.
Which is why, when you are buying AI you can not longer just buy the outcomes, you have to chose the right AI to get those outcomes. And, for all the reaons above, that may not be an LLM based product.
What I would actually do next
You may already have most of these underway, but if I were sitting on the executive team, these would be my five next steps
1. FIND IT
List how AI is being used across the business.
Include:
- AI bought by the business;
- AI embedded in existing suppliers;
- AI being developed internally; and
- shadow AI* being used unofficially.
Then identify which use cases are relatively static and which are moving feasts that could climb the autonomy spectrum as the technology changes. *AI being used unofficially by your employees
2. CLASSIFY IT
Build an autonomy register for every material AI use case.
For each one, record:
- autonomy level;
- business owner;
- purpose;
- decisions affected;
- applicable regulatory obligations;
- key controls; and
- what would trigger a reassessment.
Don't just record where the system is, record what would cause it to move.
3. STRESS IT
For every L4 or L5 use case, stress test it against:
- the regulations you fall under;
- customer outcomes;
- operational failure modes;
- supplier dependency;
- model degradation;
- human oversight; and
- concentration risk.
The question isn't "Does this work?" It's "What happens when this works differently from how we expected?"
4. PROVE IT
Take one consequential AI-enabled decision.
I'd pick the one that needs to be most defensible.
A lending decision, for example, rather than a chatbot telling someone your opening hours.
Trace it end to end:
Source data → model/system → recommendation or decision → human intervention → outcome → evidence retained
Then ask who owned each step and whether you could reproduce the decision one day or six months later. If you can't, you've found a problem worth fixing.

5. MONITOR IT
Finally, decide what would tell you that the system has:
- changed;
- degraded;
- become more autonomous;
- become unsafe; or
- started producing different customer outcomes.
Then decide:
Who owns that signal?
What threshold matters?
What happens next?
If the answer is simply "The supplier tells us" I'd keep digging.
The question I'd leave with you
The Mills Review does not create a new AI rulebook.
Instead it makes clear that the rules you already have need to be applied to systems that are becoming more autonomous, more dynamic and more deeply embedded in the decisions your organisation makes.
And that means the standard for a well-controlled AI system is increasingly simple to state:
Can you show what happened, on whose authority, and why?
If you can, you're in a much stronger position. If you can't, the problem isn't that AI regulation has changed. It's that your operating model hasn't kept up with the technology.



