# Evidence: what is tested, and how each count is produced

## EVIDENCE

- 1957 (source: `python3 scripts/published-numbers.py --metric tests`) Go-listed tests and fuzz targets, green on the self-hosted Linux CI pool for every build that ships

- 232 (source: `python3 scripts/published-numbers.py --metric break_tests`) break-tests — **not tests that pass, tests proven able to fail** when their protection is removed

- 48 (source: `python3 scripts/published-numbers.py --metric adversary_reviews`) adversarial reviews on file with a recorded verdict, by a second AI from a different vendor; 24 (source: `python3 scripts/published-numbers.py --metric sent_back`) sent the change back before it shipped. Recorded review counts are a floor.

Method: count .briefs/**/*.md files with an explicit VERDICT: SHIP, FIX-FIRST or REJECT line; exclude choice templates; one file once.

Method: count those verdict files containing at least one FIX-FIRST or REJECT verdict (even if later shipped).

### How these figures are produced

Each count on this page is generated by `scripts/published-numbers.py` from our source tree, which is not public as of September 2026.