Case Study - Financial Services

Compliance Officer AI Assistant

A mid-tier asset manager deployed an AI assistant supporting compliance officers reviewing client communications. Review throughput tripled with audit-grade source attribution and no increase in senior escalations.

The Setting

A mid-tier asset manager with around $40B in AUM, operating across the US, UK and EU. The compliance team - roughly 25 officers - reviewed client communications (emails, chat transcripts, prepared marketing material) against a regulatory framework spanning SEC, FCA, ESMA and internal policy. The framework expanded faster than the team could keep up; review backlog had grown from days to weeks, and senior compliance time was consumed clearing tier-1 reviews that could have been handled with consistent application of existing rules.

The Head of Compliance commissioned an evaluation: could an AI assistant accelerate first-pass review while preserving full audit trail, source attribution per ruling, and the existing escalation hierarchy? The non-negotiable constraint was that no AI-produced decision could stand without a named human compliance officer's sign-off; the AI was an accelerator, not a decision-maker.

What Made This Hard

High-stakes correctness

A wrong AI ruling that a compliance officer signed off on without scrutiny was the worst possible outcome. The system had to make its reasoning legible enough that an officer would notice when the AI was wrong - not just confident enough to convince them when right.

Audit-grade trail per query

Every AI ruling had to be reconstructable months later: which regulation cited, which version of which policy, which prior internal ruling, what was the LLM context window, what was the response. Standard log retention was insufficient.

Cultural resistance

Compliance officers, by professional formation, are paid to distrust assertions made without evidence. A system that produced unsourced or low-evidence answers would be ignored within a week of launch.

How the Methodology Was Applied

Total engagement duration: 18 weeks across Phases 1-3 plus Phase 4 oversight through go-live.

Phase 1 - Discovery (3 weeks)

Scope was narrowed to first-pass review of client communications against a defined subset of regulatory rules (specifically: marketing-claim review, restricted-recipient checks, and tone/promissory-language flags). Out of scope: any AI involvement in trade surveillance or anti-money-laundering - both were classified high-risk under the EU AI Act and would have required a different governance regime. Baseline metric: average reviews per officer per day, measured over the prior quarter (current state: 18 reviews/day).

Phase 2 - Architecture (4 weeks)

Retrieval-augmented generation grounded in the firm's policy library, the relevant regulatory text (SEC marketing rules, FCA conduct standards, ESMA guidelines), and three years of prior internal rulings. LLM in vendor private tenant (Azure OpenAI) under firm-managed deployment to satisfy data-residency requirements per region. Response schema required structured output: ruling (pass/flag/escalate), confidence band, source citations with pinpoint paragraph references, reasoning summary in compliance-officer vocabulary.

Phase 3 - Governance Design (2 weeks, parallel to Phase 2)

Twelve-control baseline tailored to regulatory audit needs. The dedicated control of note: per-query immutable log capturing the full context window, the model version, the response, the officer's accept/override decision, and timestamps. Logs were stored in a separate audit log store with 7-year retention matching the firm's compliance retention policy. Risk classification under EU AI Act: limited-risk (transparency obligations applied, but not the full high-risk regime). Named accountable owner: Head of Compliance, with the Chief Risk Officer as escalation.

Phase 4 - Implementation Oversight (12 weeks)

Engineering was split: the AI components (retrieval, LLM orchestration, response schema enforcement) were built by a specialist delivery partner; the integration into the existing compliance workflow tool was built by the firm's internal RegTech team. Three governance gates were enforced: gate 1 at week 4 (data and access controls verified), gate 2 at week 8 (audit log completeness and immutability), gate 3 at week 11 (calibration of confidence thresholds against a golden set of historical rulings).

Measured Against the Baseline

Review throughput

3.1x increase per officer per day (from 18 baseline to median 56 reviews/day on AI-assisted workflow). Backlog cleared within 7 weeks of pilot launch.

Escalation rate to senior compliance

Flat against baseline. Critically: the AI did not increase escalations, proving it was applying rules consistently with how senior compliance had been applying them - the most important quality signal in this domain.

Source attribution

100% of rulings shipped with at least one pinpoint citation. Average 2.3 citations per ruling. Audit spot-check accuracy at 98%.

Officer override rate

11% in steady state. Within the 8-15% target band - high enough to prove officers were exercising judgement, low enough to demonstrate value above pure manual review.

Audit readiness

Sample-pull tested in 60 seconds. Regulator sample queries returned a complete reconstructable record per query, including model version and context-window snapshot.

Adoption

23 of 25 officers active daily within 8 weeks of launch. The two non-adopters were transitioning out of role for unrelated reasons; no compliance officer who remained in role declined to use the tool.

What We Would Change

The golden set should be larger from the start

The Phase 3 calibration of confidence thresholds used a 200-ruling golden set assembled from historical decisions. Two months in, the team realized the golden set under-represented edge cases that occur rarely but matter when they do (specific cross-jurisdictional marketing scenarios). Extending the golden set should have been a Phase 1 deliverable - bigger upfront, not added later.

Override-reason capture from day one

When officers overrode the AI, the initial UI captured the override decision but not the reason. Introducing structured override-reason capture in month 3 would have allowed earlier model and prompt refinement. Future engagements: override-reason capture is part of the launch UX, not an add-on.

Regulator pre-engagement worth the effort

The firm voluntarily briefed the FCA on the system architecture before deployment. That briefing turned out to be material when the broader RegTech market drew regulator scrutiny six months later; the firm's prior engagement was treated as evidence of good-faith implementation. Future engagements in regulated industries: we now recommend voluntary regulator pre-engagement as a Phase 3 deliverable.

Other Engagements

Healthcare - HIPAA-Compliant Clinical Knowledge RAG

Regional healthcare provider, clinical-knowledge assistant, 42% reduction in time-to-answer with zero PHI leakage.

Read this case study

Manufacturing - Document AI

European industrial manufacturer, supplier QA document extraction, 78% auto-processed end-to-end.

Read this case study

AI Governance

The maturity model and twelve-control baseline this case study applied.

Read the AI Governance page

A Financial-Services AI Initiative in Mind?

An Architecture Review identifies the regulatory tier, the highest-risk architecture decision, and whether the initiative is worth taking past Discovery.