Evidence: what is tested, and how each count is produced
EVIDENCE
- 1957 Go-listed tests and fuzz targets, green on the self-hosted Linux CI pool for every build that ships
- 232 break-tests — not tests that pass, tests proven able to fail when their protection is removed
- 48 adversarial reviews on file with a recorded verdict, by a second AI from a different vendor; 24 sent the change back before it shipped. Recorded review counts are a floor.
Method: count .briefs/**/*.md files with an explicit VERDICT: SHIP, FIX-FIRST or REJECT line; exclude choice templates; one file once.
Method: count those verdict files containing at least one FIX-FIRST or REJECT verdict (even if later shipped).
How these figures are produced
Each count on this page is generated by scripts/published-numbers.py from our source tree, which is not public as of September 2026.