Corpus PII-Prevalence Audit: Sample Evidence Pack Synthetic sample · no real PII

Calibrated PII-prevalence audit · 95% confidence interval · generated 2026-07-03 12:43 UTC

Calibrated corpus PII prevalence
21.53%
95% CI  [16.48%, 26.59%]
Corpus
3,000
Sample judged
450
Human gold
160
Raw flag-rate
28.89%
Excluded (UNCALIBRATED)
1

The headline is a stratified prediction-powered (PPI++) estimate: per stratum it anchors on the human gold labels and uses the judge's score as a variance-reduction signal, then design-weights the strata to the corpus. It is a two-sided misclassification correction (it corrects the judge's misses and its false alarms), not a raw count. It is computed over the calibrated and pooled strata only; ill-posed strata are reported separately and never fold into the headline.

1Per-stratum prevalence

StratumStatusPrevalence95% CI SampleFlaggedGold PII/cleanWeight
support-ticketsCALIBRATED32.59%[24.99%, 40.20%]2487028/600.550
wiki-pagesPOOLED3.10%[0.00%, 7.68%]148142/510.330
id-verification-logsUNCALIBRATED83.93%[73.34%, 92.23%]544619/00.120

CALIBRATED the PPI++ rectifier is anchored on this stratum's own gold subset.   POOLED gold below the 5-each-class floor; the rectifier rests on very few labels.   UNCALIBRATED calibration ill-posed (a gold class is empty); raw flag-rate only, excluded from the headline.

2Honesty caveats

POOLEDwiki-pages: insufficient gold (have 2 PII / 51 clean, need 5/5) -> reported POOLED (the PPI rectifier rests on very few labels); label 3 more PII-positive to calibrate it locally.
UNCALIBRATEDid-verification-logs: insufficient gold (have 19 PII / 0 clean, need 5/5) -> reported UNCALIBRATED (raw flag-rate, excluded from the headline); label 5 more clean to calibrate it.
Dev vs validatedThe judge prompt and taxonomy are the development version; the estimator (stratified PPI++ with Monte-Carlo uncertainty propagation) is the validated build. This bound is only as strong as the human gold it rests on, it is a calibrated estimate from a sample, not a census or a scanner.

3Method

4Provenance · reproducibility

corpus audit_cli/sample_evidence_pack/synthetic_corpus.jsonl
strata field:source
sample seed=23 method=neyman n=450
gold seed=41 subset=160
calibrate seed=31 confidence=0.95
note SYNTHETIC corpus, deterministic, no real PII