AI & Compliance
Transaction Monitoring Is Not in Annex III
As of today, every high-risk AI system in production has to meet the EU AI Act's heaviest obligations. Look at what Annex III actually names on the financial services side, then look at what the FCA has spent five years fining banks for, and the two lists barely overlap.
By Thomas Geater· 3 August 2026 · 8 min read
Today is the EU AI Act’s high-risk deadline. Every Annex III system already in production had to be compliant by today regardless of when it was built or procured, with penalties running to €15 million or 3% of global turnover. Today is the point where the Act stops being something firms prepare for and becomes something they are measured against.
Which makes it worth reading Annex III slowly, because on the financial services side it is a great deal shorter than most people assume.
It names AI systems used to evaluate the creditworthiness of natural persons or establish their credit score. It names AI systems used for risk assessment and pricing in life and health insurance. That is the list. Transaction monitoring is not on it. Sanctions and watchlist screening is not on it. Customer risk rating is not on it. The systems sitting at the centre of every financial crime function in Europe are, as a category, absent.
Now set the enforcement record beside it. Since 2021 the FCA has issued thirteen fines totalling roughly £300.7 million against firms for anti-money-laundering systems and controls failings. Nationwide was fined £44 million, and among the findings was that transaction monitoring alert thresholds had been set very high. Monzo, £21.1 million. Barclays, £42 million. Across 2026 the running total for systems and controls failings generally goes past £316 million.
Whatever else is true, that is where the regulatory pain in financial crime actually lands. And it lands on automated decision-making almost every time.
The mechanism, stated plainly
The automated systems whose failures regulators punish sit outside the regime built to govern high-risk AI. Credit scoring — consequential, yes, but comparatively well understood and already heavily supervised — sits inside it.
That is not a criticism of the drafters. Annex III was built on a coherent theory of harm: systems determining access to essential services, employment, education and credit, where an opaque model can quietly shut somebody out of ordinary economic life. Credit scoring fits the theory exactly.
The oddity is that transaction monitoring fits it too, and nobody put it on the list.
Consider what a monitoring system does to a person. It scores their behaviour against a model of what suspicious looks like. When it fires, the consequence is not a declined application they can appeal and reapply for. It is a frozen account, a payment held, a customer exit executed without explanation, and in the more serious cases a suspicious activity report submitted to a financial intelligence unit that the subject is never told about and cannot see. Tipping-off provisions mean the affected person usually cannot be given a reason, cannot contest the assessment, and cannot correct the data underneath it. Measured against almost any definition of consequential automated decision-making, that sits at the sharp end.
And the calibration of the whole thing is a human choice, made in a room, by somebody.
The Nationwide finding illustrates it better than any hypothetical. The failure was not a model that malfunctioned. It was a threshold set at a level that suppressed the alerts that should have been investigated. Nobody has to write a bad algorithm to produce that outcome. Somebody has to set a number.
Why “it’s not AI, it’s rules” doesn’t rescue this
The standard objection is that most transaction monitoring isn’t AI at all. Much of it is deterministic rules, scenario libraries and threshold logic predating the current wave of machine learning by two decades, and the Act governs AI systems rather than spreadsheets.
Weakening every quarter, that objection, and it was never as strong as it sounded.
Firms across the sector are layering machine learning over rules-based estates specifically to cut false positives, introducing anomaly detection, behavioural clustering and alert-scoring models that sit between the rule firing and the analyst reading it. The moment a model decides which alerts a human sees, the human oversight question is live, whatever the scenario library underneath is doing. An alert-triage model suppressing ninety per cent of what a rules engine produces is making the same kind of decision the threshold was making, with less visibility into how.
There is a subtler version of the point too. The Act’s obligations for high-risk systems — Article 9 risk management, Article 14 human oversight designed in rather than bolted on, technical documentation good enough for an assessor to trace how a decision got produced — are not exotic AI-specific inventions. They are, almost line for line, what competent model risk management already asks for. The controls the Act forces onto a credit scoring model are the controls a monitoring threshold decision needs. In the AML estate, nothing compels them.
What actually governs the alert threshold
Less than most people think, and nothing that produces artefacts an assessor could inspect in advance.
In the UK it is the FCA’s outcomes-based supervision, model risk expectations, and the firm’s own governance. Which is to say it is enforced retrospectively, after a failure, through a penalty notice. A real mechanism, and £300 million says it has teeth. But it works the way a coroner works: it tells you what went wrong once it has already gone wrong, at an institution large enough for the outcome to become visible. It requires nobody to document, before deployment, why a threshold sits where it sits and who answers if it is wrong.
The FCA has been explicit that it does not intend to write AI-specific rules, preferring existing principles and the Senior Managers regime. There is a genuine argument for that and I have written elsewhere about why principles-based supervision may age better than prescriptive rules. It does mean that in the UK the alert threshold is governed by principle and by hindsight. In the EU, for these particular systems, it is not governed by the AI Act at all.
Two regulatory philosophies arriving at the same gap from opposite directions. Worth noticing.
The practical question
If your firm ran an AI Act scoping exercise this year, there is a fair chance it worked the way most scoping exercises work. Map the estate against Annex III, identify what lands inside, build the compliance programme around those systems, record everything else as out of scope. Done that way the monitoring and screening estate comes back clean, because Annex III does not name it.
Clean and unmapped are different things.
So the question worth asking on Monday is not whether transaction monitoring is in scope for the AI Act. It very likely is not. The question is whether anyone can currently produce, for the alert-scoring and triage models running in production right now, the four things Article 14 would have demanded if they were: a written statement of what the system is for, a record of who set the current calibration and on what basis, evidence that a human reviewing an output has enough information to overrule it, and a traceable account of how any given alert was generated.
If those exist, the Act’s absence costs nothing. The discipline is already there.
If they don’t, then the most consequential automated decisions the firm makes about its customers are being made by systems no regime currently requires anybody to explain — and the record suggests that gets discovered, on average, about £23 million at a time.
Written by
Thomas Geater
I write The Typology Files, on financial-crime compliance and what generative and agentic AI are doing to it. BSc Business Management with Finance, University of Brighton. I work independently and take on advisory and writing engagements.