Reproducible Audit

We Audited AI Chatbots on Banking Regulation

3 chatbots. 15 questions a bank risk officer would ask. Every answer verified against the actual regulation. The AI gets the risk right — but the rules wrong.

0%
Confidence score on EU AI Act
(Microsoft Copilot)
15
Unique regulatory questions
25 verified model responses
Risk framework identified
Rules failed

What We Found

🎯

Risk Identification ≠ Regulatory Compliance

All three chatbots correctly identified regulatory concepts and risk categories. But when verified against actual articles, citations were wrong, outdated, or fabricated. A non-specialist would never notice.

📋

Draft vs. Final Regulation Confusion

ChatGPT cited article numbers from the EU AI Act proposal draft, not the final regulation (2024/1689). The numbering changed significantly — Copilot used numbering from neither version.

⚠️

Fabricated Regulatory References

Multiple responses cited or implied regulatory references that were missing, outdated, or did not support the claim. Confident tone, wrong substance.

🔍

SR 26-02 "GenAI Gap" Confirmed

The new US guidance (SR 26-02, April 2026) explicitly excludes GenAI from model risk scope. Our audit shows why: these models can't reliably cite the rules they claim to follow.

Chatbot Performance

ChatGPT (GPT-4o)
Mixed
Best at structure. Used draft regulation numbering. Fabricated specific ESMA guidelines.
Google Gemini
Mixed
Most cautious framing. Still cited non-existent articles. Missed critical exemptions.
Microsoft Copilot
Weakest
0% PLV confidence on EU AI Act. Used numbering from neither draft nor final regulation.

Why This Matters Now

EU AI Act SR 26-02 DORA

EU AI Act high-risk obligations (Annex III) apply from December 2, 2027 under the May 2026 Omnibus Agreement. Articles 9–15 set core high-risk system obligations, including risk management, data governance, technical documentation, logging, transparency, human oversight, accuracy, and robustness. Article 27 adds fundamental-rights impact assessment obligations for many deployers. The postponement from August 2026 creates a preparation window — not a pause. Banks deploying AI chatbots for regulatory advice, customer service, or risk assessment need to demonstrate that outputs are accurate and auditable.

SR 26-02 (US, April 2026) replaced 15 years of model risk guidance — and explicitly excluded GenAI/agentic AI from scope, calling it "novel and rapidly evolving." 88% of financial institutions already have AI in production.

DORA raises the bar for ICT risk management, resilience, and auditability in financial entities. AI-assisted regulatory workflows need evidence trails that can survive operational and supervisory review.

Verify Your AI Before Regulators Do

PLV (Plan-Level Verification) automatically checks whether AI outputs faithfully follow regulatory procedures — article by article, step by step.

Request Evidence Pack Read Full Methodology