The headline is a stratified prediction-powered (PPI++) estimate:
per stratum it anchors on the human gold labels and uses the judge's score as a
variance-reduction signal, then design-weights the strata to the corpus. It is a two-sided
misclassification correction (it corrects the judge's misses and its false alarms), not a
raw count. It is computed over the calibrated and pooled strata only; ill-posed strata are
reported separately and never fold into the headline.
1Per-stratum prevalence
Stratum
Status
Prevalence
95% CI
Sample
Flagged
Gold PII/clean
Weight
support-tickets
CALIBRATED
32.59%
[24.99%, 40.20%]
248
70
28/60
0.550
wiki-pages
POOLED
3.10%
[0.00%, 7.68%]
148
14
2/51
0.330
id-verification-logs
UNCALIBRATED
83.93%
[73.34%, 92.23%]
54
46
19/0
0.120
CALIBRATED the PPI++ rectifier is anchored on this stratum's own gold subset.
POOLED gold below the 5-each-class floor; the rectifier rests on very few labels.
UNCALIBRATED calibration ill-posed (a gold class is empty); raw flag-rate only, excluded from the headline.
2Honesty caveats
POOLEDwiki-pages: insufficient gold (have 2 PII / 51 clean, need 5/5) -> reported POOLED (the PPI rectifier rests on very few labels); label 3 more PII-positive to calibrate it locally.
UNCALIBRATEDid-verification-logs: insufficient gold (have 19 PII / 0 clean, need 5/5) -> reported UNCALIBRATED (raw flag-rate, excluded from the headline); label 5 more clean to calibrate it.
Dev vs validatedThe judge prompt and taxonomy are the development version; the estimator (stratified PPI++ with Monte-Carlo uncertainty propagation) is the validated build. This bound is only as strong as the human gold it rests on, it is a calibrated estimate from a sample, not a census or a scanner.
3Method
Sampling Stratified neyman allocation, n=450, drawn without replacement from a seeded RNG.
Calibration Per-stratum stratified PPI++ (prediction-powered inference): anchors on the gold labels and uses the judge's score as a variance-reduction covariate; thin strata are flagged POOLED (the rectifier rests on very few labels); empty-gold-class strata reported uncalibrated and excluded (UNCALIBRATED). Corpus CI via the design-weighted PPI++ combine.
Judgemock (local / loopback - no data egress).
Gold budget 160 human-labelled documents, drawn stratified-random from the judged sample (exchangeable with the remainder).
4Provenance · reproducibility
corpus audit_cli/sample_evidence_pack/synthetic_corpus.jsonl