Evidence
Validation evidence (scoped claims)
This page supports the site’s scoped empirical language (for example “0 false allows observed in a 120-case validation set (Standard tier)”). It is not a claim of zero false allows in production forever.
Public packaging in progress.
Case files and run summaries currently live in private product repos on the operator machine.
The next packaging step is to publish a read-only bundle (cases + summary JSON + method note) and link it here.
Until that ships, treat numbers on marketing pages as internally backed, not yet visitor-recomputable.
120-case set (PLV / cascade lineage)
- Cases:
pot-cli/cases/plv-120-cases-v3.json(n=120) - Canonical runs dir:
thoughtproof-api-v2/runs/canonical-120/ - Example (2026-05-24 summaries): multiple cascade configs reported
false_allows: 0on n=120 - Counter-example (do not hide): a later combined-nano-solo summary reported
false_allows: 1— config and date matter - Method notes:
docs/plv-benchmark-run-comparison.md
Other false-allow observations (different suites)
- DQL stop-case FAR: live report with FAR 0.00 on a 21-scenario stop-case set (ADSB / dql-benchmark lineage) — not the PLV 120-set
- Bakeoff tables: small-n method proofs with explicit CI caveats
What we will not claim
- Unscoped “0 false ALLOWs” across all products and time
- General “98.1% accuracy” without a named metric and set
- That every future run on the 120-set stays at zero FA