AI & Compliance
Agentic AI in Compliance: Who's Accountable When the Analyst Is a System
Agentic AI can already run adverse media research, draft EDD narratives, and triage alerts with no human touching the middle steps. The technology changes fast. The accountability question doesn't move at all — and most teams haven't worked out where it actually sits.
July 28, 2026 · 10 min read
There’s a meaningful difference between a chatbot that answers a compliance question and an agentic system that goes and does the compliance work — pulls the adverse media, queries the corporate registry, drafts the EDD write-up, flags the file for review — across several steps, with no person in the loop until the end. That second kind is already in production at a growing number of institutions, not as a pilot but as the first pass on real case volume. The capability is moving fast. The question of who’s actually accountable when it gets something wrong hasn’t moved at all, and it’s worth being precise about why it can’t.
Agentic doesn’t mean autonomous in the sense that matters
“Agentic” describes how the system executes — multiple steps, tool use, limited human intervention mid-task. It says nothing about who’s answerable for the output, because in UK financial services that question was never a function of who (or what) did the work. Under the Senior Managers and Certification Regime, every prescribed responsibility has to sit with a named senior manager, full stop. Deploying an agentic system inside that function doesn’t create a gap in the org chart for the AI to fill — the senior manager who owns customer due diligence still owns it, whether the first-draft research was produced by a graduate analyst or a language model chaining together six tool calls. The system can execute the task. It cannot hold the responsibility, because the entire regulatory architecture was built around a named accountable person, not a process.
This is worth stating plainly because a lot of vendor material implicitly suggests otherwise — that sufficiently capable agentic tooling somehow absorbs risk rather than relocating where the evidence trail needs to point. It doesn’t. It just changes what the senior manager needs to be able to demonstrate they were overseeing.
Where the FCA’s approach actually lands on this
The FCA has been explicit, including from its chief executive as recently as December 2025, that it isn’t introducing AI-specific rules — the stated reasoning is that the technology changes on a three-to-six-month cycle, too fast for bespoke rules to stay current, and existing frameworks (SMCR, Consumer Duty, model risk management expectations) already cover the ground that matters. The FCA’s AI Lab and its AI Live Testing pilot exist to let firms test deployments collaboratively with the regulator rather than to hand down a new AI rulebook.
The practical reading for a compliance function: there is no separate “AI compliance checklist” waiting to be published that will tell you how to govern an agentic EDD tool. The governance obligations that already applied — can you show the output was overseen, can you show a person made the judgment call, can you show the process is periodically reviewed for whether it’s actually working — apply to the agentic version of the task exactly as they applied to the manual version. Firms waiting for AI-specific guidance before they formalize oversight of their agentic tools are waiting for something the regulator has said, twice now, isn’t coming.
The failure mode that’s genuinely new
None of that means nothing has changed. Agentic systems introduce a specific new risk that a human researcher doesn’t carry in the same way: fluent, confident, and wrong. A junior analyst who isn’t sure of a fact tends to hedge, flag it, or leave it out. A language-model-driven agent asked to summarize a source will produce a summary in the same confident register whether the underlying source actually says what the summary claims or not. The failure doesn’t look like a gap — it looks like a finished piece of work.
That changes what oversight has to check for. Reviewing agentic output for completeness isn’t enough; someone has to spot-check the underlying claims against the actual sources, on a sampling basis proportionate to risk, precisely because the failure mode is designed (by the nature of how these models generate text) to be unobtrusive. A due diligence file that looks thorough and reads well is not evidence that it’s accurate — it’s only ever been evidence that it’s fluent, and agentic tools are exceptionally good at fluent.
What this means for a compliance function building toward this now
Three things worth doing before an agentic tool touches live casework rather than after: name the accountable senior manager for that specific function explicitly, in writing, before deployment, not as an afterthought once something goes wrong. Build the sampling-based verification step into the process design from day one, sized to the risk level of the task, rather than trusting output quality as a proxy for output accuracy. And keep a clear, demonstrable line between what the agent drafted and what a human reviewed and approved — not because the technology is unreliable, but because in a regime built entirely around named human accountability, that line is the only thing that was ever actually being tested.