Full methodology, questions, model responses, regulatory references, and Plan-Level Verification (PLV) verdicts. Prepared for N26 risk, compliance, and product leadership.
Key Finding: All three chatbots correctly identified regulatory risk concepts, but failed to deliver regulatorily compliant answers. Citation errors, draft vs. final regulation confusion, and missing procedural steps were systemic.
Implication for N26: AI chatbots are useful for initial research on established regulation but unreliable for emerging rules (EU AI Act, DORA, SR 26-02). Every regulatory output requires independent verification before use in compliance documentation or supervisory submissions.
15 questions across five regulatory domains relevant to banking risk officers:
Every response verified with Plan-Level Verification (PLV) against primary regulatory sources (Regulation 2024/1689, BCBS d457, MiFID II Directive 2014/65/EU, etc.).
Verdict tiers: ALLOW (faithful), UNCERTAIN (partial), BLOCK (material gaps).
| Model | Responses | ALLOW | UNCERTAIN | BLOCK | Avg. PLV Confidence |
|---|---|---|---|---|---|
| ChatGPT (GPT-4o) | 15 | 8 (53%) | 5 (33%) | 2 (13%) | 52% |
| Google Gemini | 5 | 1 (20%) | 3 (60%) | 1 (20%) | 34% |
| Microsoft Copilot | 5 | 0 (0%) | 2 (40%) | 3 (60%) | 33% |
Performance gap by regulation age: 100% ALLOW on 15-year-old SR 11-7 vs 0% ALLOW on EU AI Act / DORA (finalized 2024–2025). This pattern is critical: banks need AI most for new regulation, where it is least reliable.
Question: "We're building an ML-based credit scoring system for consumer loans. Under the EU AI Act, is this automatically classified as high-risk, and what obligations does that trigger?"
"Yes — under the EU AI Act, an ML-based credit scoring system for consumer loans is automatically classified as a high-risk AI system."
Correctly cited Annex III Category 5(b) and listed Chapter 2 obligations. Missed fraud detection exemption and Article 27. Used proposal-era chapter numbering (Chapter 2 instead of final Chapter 3, Section 2).
Question: "We're launching an automated investment advisory service for retail clients in the EU. Does MiFID II treat this differently from human advisory, and what specific requirements apply?"
Correctly stated MiFID II is technology-neutral and described suitability, product governance, disclosure, and record-keeping. Referenced "ESMA guidance" at high level.
Question: "Our bank wants to use the Internal Models Approach for market risk capital. Under the finalized Basel III reforms (FRTB), can we still use our existing VaR models?"
Strong technical description of VaR → Expected Shortfall transition, desk-level approval, P&L attribution, backtesting, NMRF, and Default Risk Charge.
Pattern across remaining 12 questions: Similar issues — correct high-level risk identification with missing or incorrect article numbers, absent exemptions, and failure to cite the exact regulatory text a compliance officer would need for implementation. Full raw responses and PLV traces available on request.
Plan-Level Verification evaluates each response against a structured regulatory checklist derived from primary sources. Each verification step receives a 0.00–1.00 score. Overall confidence is the weighted average of step scores.
| Verdict | Confidence Range | Meaning |
|---|---|---|
| ALLOW | ≥ 80% | Response faithfully follows regulatory procedure; suitable for compliance use with minimal additional review. |
| UNCERTAIN | 20–79% | Directionally correct but missing critical procedural elements, citations, or exemptions. Requires human verification. |
| BLOCK | < 20% | Material gaps that would create compliance risk if used without correction. Not suitable for regulatory documentation. |