← Back to Audit

Banking AI Chatbot Audit — Evidence Pack

Conducted May 2026 by ThoughtProof • 15 questions • 25 verified responses • PLV verification

Full methodology, questions, model responses, regulatory references, and Plan-Level Verification (PLV) verdicts. Prepared for N26 risk, compliance, and product leadership.

Executive Summary

Key Finding: All three chatbots correctly identified regulatory risk concepts, but failed to deliver regulatorily compliant answers. Citation errors, draft vs. final regulation confusion, and missing procedural steps were systemic.

0%
Copilot confidence on EU AI Act (E-02)
53%
ChatGPT ALLOW rate across 15 questions
2
BLOCK verdicts (high compliance risk)
100%
ChatGPT on legacy SR 11-7 (now superseded)

Implication for N26: AI chatbots are useful for initial research on established regulation but unreliable for emerging rules (EU AI Act, DORA, SR 26-02). Every regulatory output requires independent verification before use in compliance documentation or supervisory submissions.

Methodology

Questions & Categories

15 questions across five regulatory domains relevant to banking risk officers:

  • A. Model Validation — SR 11-7 / SR 26-02 (3 questions)
  • B. Capital & Risk — Basel III/FRTB, CRR3 (3 questions)
  • C. AI/ML Governance — BCBS 239, EBA (3 questions)
  • D. Compliance Edge Cases — MiFID II, AML/BSA (3 questions)
  • E. Emerging Regulation — EU AI Act, DORA (3 questions)

Models & Verification

  • ChatGPT (GPT-4o) — answered all 15
  • Google Gemini — 5 priority questions
  • Microsoft Copilot — 5 priority questions

Every response verified with Plan-Level Verification (PLV) against primary regulatory sources (Regulation 2024/1689, BCBS d457, MiFID II Directive 2014/65/EU, etc.).
Verdict tiers: ALLOW (faithful), UNCERTAIN (partial), BLOCK (material gaps).

Results Summary

Model Responses ALLOW UNCERTAIN BLOCK Avg. PLV Confidence
ChatGPT (GPT-4o) 15 8 (53%) 5 (33%) 2 (13%) 52%
Google Gemini 5 1 (20%) 3 (60%) 1 (20%) 34%
Microsoft Copilot 5 0 (0%) 2 (40%) 3 (60%) 33%

Performance gap by regulation age: 100% ALLOW on 15-year-old SR 11-7 vs 0% ALLOW on EU AI Act / DORA (finalized 2024–2025). This pattern is critical: banks need AI most for new regulation, where it is least reliable.

Detailed Question Breakdown

E-02 — Credit Scoring under EU AI Act (High-Risk Classification)

Question: "We're building an ML-based credit scoring system for consumer loans. Under the EU AI Act, is this automatically classified as high-risk, and what obligations does that trigger?"

Microsoft Copilot

"Yes — under the EU AI Act, an ML-based credit scoring system for consumer loans is automatically classified as a high-risk AI system."
Sources cited: EBA summary page, third-party blog. No articles from Regulation 2024/1689.
PLV Verdict: BLOCK (0% confidence) — Failed all 6 verification steps. Missed Annex III 5(b), Articles 9–15, fraud exemption, Article 27 fundamental rights impact assessment, and August 2026 application date.

ChatGPT (GPT-4o)

Correctly cited Annex III Category 5(b) and listed Chapter 2 obligations. Missed fraud detection exemption and Article 27. Used proposal-era chapter numbering (Chapter 2 instead of final Chapter 3, Section 2).
PLV Verdict: UNCERTAIN (12% confidence)
Gold Standard Reference: Regulation (EU) 2024/1689, Annex III point 5(b), Articles 9–15, Article 27, Article 113(3) application date, Article 111(2) grandfathering.
D-03 — Robo-Advisory under MiFID II

Question: "We're launching an automated investment advisory service for retail clients in the EU. Does MiFID II treat this differently from human advisory, and what specific requirements apply?"

ChatGPT (GPT-4o)

Correctly stated MiFID II is technology-neutral and described suitability, product governance, disclosure, and record-keeping. Referenced "ESMA guidance" at high level.
PLV Verdict: BLOCK (12% confidence) — Failed on Article 25(2), Delegated Regulation (EU) 2017/565 Articles 54–56, and ESMA35-43-3172 suitability guidelines. No specific citations.
Gold Standard: MiFID II Article 25(2), Delegated Regulation (EU) 2017/565 Arts. 54–56, ESMA Guidelines ESMA35-43-3172 (Sept 2022).
B-01 — FRTB Internal Models Approach (VaR vs Expected Shortfall)

Question: "Our bank wants to use the Internal Models Approach for market risk capital. Under the finalized Basel III reforms (FRTB), can we still use our existing VaR models?"

ChatGPT (GPT-4o)

Strong technical description of VaR → Expected Shortfall transition, desk-level approval, P&L attribution, backtesting, NMRF, and Default Risk Charge.
PLV Verdict: UNCERTAIN (45% confidence) — Did not distinguish original Basel III (2010) from finalized reforms (2017). Missed CRR3 (Regulation 2024/1623) and BCBS d457 as primary sources. Jurisdictional implementation gaps for EU vs US.
Gold Standard: BCBS d457 (FRTB), Regulation (EU) 2024/1623 (CRR3), Basel III Endgame (US implementation).

Pattern across remaining 12 questions: Similar issues — correct high-level risk identification with missing or incorrect article numbers, absent exemptions, and failure to cite the exact regulatory text a compliance officer would need for implementation. Full raw responses and PLV traces available on request.

PLV Confidence Scoring Methodology

Plan-Level Verification evaluates each response against a structured regulatory checklist derived from primary sources. Each verification step receives a 0.00–1.00 score. Overall confidence is the weighted average of step scores.

VerdictConfidence RangeMeaning
ALLOW≥ 80%Response faithfully follows regulatory procedure; suitable for compliance use with minimal additional review.
UNCERTAIN20–79%Directionally correct but missing critical procedural elements, citations, or exemptions. Requires human verification.
BLOCK< 20%Material gaps that would create compliance risk if used without correction. Not suitable for regulatory documentation.
Full raw model responses, complete question list (A-01 to E-03), and PLV verification traces available on request.
Contact: raul@thoughtproof.ai • Audit conducted May 2026 using PLV thorough_balanced tier.