prevalence-audit

Product

How much personal information is hiding in your data? Get a number you can prove.

Big datasets are impossible to read by hand, so most companies genuinely do not know how much personal information they are holding: names, emails, addresses, health details, and the like. When a regulator, an auditor, or a customer asks, you need a real number, not a guess. We measure it from a sample, check that sample against your own reviewers, and give you a rate with an honest margin of error. A number you can put on the record.

Runs on your own computers, nothing is sent out One number for the whole dataset, not a pile of alerts Checked against your own reviewers

Tested on public datasets where the true answer was already known. Our number landed on target every time.

Pricing

From open core to run-it-yourself to done-for-you.

Start free with the open-core tool, license it and run the audit yourself, or have us certify one dataset or stand behind all of your data. Platforms and large deployments have their own path. You get the same evidence pack throughout.

Open source · MIT · $0

The drift-aware core, free and open.

The estimator that produces the calibrated rate is MIT-licensed and free to read, run, and build on. It is releasing with our preprint. Add your email and we will tell you the day it is public.

Get notified at release
License · run it yourself

Local tool license

$49/yr
per person; teams from $3,000/yr

Run the audit yourself, on your own computers.

  • The same evidence pack, produced in-house
  • Runs on your own computers, so no document leaves your building
  • Personal license $49/yr; teams from $3,000/yr
Get a license
Evidence pack · fixed scope

Compliance Evidence Pack

On request
fixed scope, one dataset

The evidence you have been asked to produce, that stands up.

For a privacy regulator's inquiry, a lawsuit, a merger or a breach, or an EU AI Act training-data summary. We certify a single dataset and hand you the pack your legal team can put on the record.

  • A measured rate, with a 95% margin of error
  • A breakdown by source
  • A written method anyone can re-check
Request an evidence pack
Diagnostic · we run all your data

Prevalence Diagnostic

On request
one-time, sized to your data

We run all of your data and stand behind the number.

  • Everything in the evidence pack, across every source
  • We size the dataset and the sampling plan with you
  • We stand behind the result with your auditors and lawyers

Keep it current: a monitoring retainer, scoped with the engagement. We re-run it on a schedule and update the number as your data changes.

Scope a diagnostic
Enterprise / OEM · pricing on request

Embed the audit, or run it across many teams.

For platforms embedding the audit, such as virtual data rooms and data-room providers, and for large multi-team deployments. Priced to your deal volume and estate, not off a list. We scope it with you.

Contact sales

Audit → gate: keep the number true, in CI

The one-time audit is also the calibration bootstrap for a drift-aware gate.

The audit produces exactly the artifacts a recurring check needs: per-source detector accuracy, a reference score distribution, and a gold subset. Point that gate at your build pipeline and it re-checks the same number on every release, on your own machines, and tells you the moment it goes stale — with a priced quote for exactly what it costs to fix, never a bare failure. Every price below is computed from this repo's own sizing model, not a round number picked to look right; the full derivation and the script that produced it are in demo/PRICING-DERIVATION.md.

The Audit SKU below is a self-serve, sized-by-the-planner version of the same audit as the Compliance Evidence Pack and Prevalence Diagnostic above — pick it when your target margin of error and rough corpus size are already known; scope a custom engagement above when they aren't, or when you need us to stand behind the result with your auditors.

Audit SKU · available now

Prevalence audit, gate-ready

$1,260–$2,200
one-time, sized to your target margin of error

The same audit above, sized and priced by the sampling planner itself.

TierTarget CIMin. corpusGold labelsPrice
Starter±3.0pp1,250 docs126$1,260
Standard±2.5pp1,800 docs166$1,580
Precision±2.0pp2,750 docs242$2,200

Computed at a 92%/97% sensitivity/specificity operating point over a 4% base rate (this repo's own sizing defaults); a real engagement re-derives every number from your judge's measured accuracy and a small pilot. Margins of error in this ±2-3pp band are designed and simulation-verified at these operating points, not a guarantee reproduced on your data. Below the listed minimum corpus for a tier, the planner refuses to promise the width — e.g. a 500-document corpus cannot reach even ±3.3pp at this operating point, and the audit says so instead of quietly missing.

Scope a gate-ready audit
Gate subscription · early access / roadmap

CI/CD drift gate

$24/run compute floor
per-run metered; no committed list price yet

Not shipping for general availability yet — this is the priced floor for the waitlist.

  • Samples ~1,200 docs per run, judged on your own machines via the code-enforced egress guard (EgressBlockedError); one open item on in-loop evaluation paths is disclosed in our DPA checklist
  • Compute floor = 1,200 docs × $0.02/doc = $24.00/run; an indicative early-access range is $48–$96/run at a 2–4x margin — not a final price
  • WATCH is the default posture: a stale-looking signal never fails your build by itself. A hard STALE stop is opt-in, per corpus, once your team has tuned it
  • False-stop rate is measured per corpus with a Wilson 95% upper-bound (e.g. ≤25.5% on our own demo corpus at un-tuned thresholds) — never a universal number; every engagement measures and tunes its own
  • No hard SLA yet. Best first fit: short-document pipelines (tickets, logs, config dumps) — our own synthetic support/devops corpora demonstrate the judge's behavior at this cadence, not a detection claim on your real tickets
Join early access
Recalibration packs · early access

When the gate says STALE

$810–$1,560
per event, metered — not purchasable standalone yet

"Blocked; $X of relabeling un-blocks it" — the gate prints the quote, not just a failure.

TierStandard resample
Starter$810
Standard$1,070
Precision$1,560

Priced on the same math gate will run once a real judge is wired — not purchasable standalone yet. Today's gate runs only on an offline mock judge, so a real STALE trigger can't occur yet; the quote above is the exact sizing math the gate will print on a real STALE verdict. Active sampling can reduce the gold-label bill further — quantified per corpus at audit time, not priced as a standing discount.

Ask about recalibration packs

privacy-judge gate is a prototype: both commands run against an offline synthetic corpus with a mock judge today (patent pending, US Provisional Application No. 64/101,018, on the drift-sentinel method underneath it). Wiring a real local judge is the next integration step and does not change the baseline schema, the exit-code contract, or any number on this page.

What you can measure

Personal information is just the first thing you can count this way.

If a trained person can label it, we can put a number on how often it turns up in a pile of documents too big to read by hand. Below: what we have already proven, and where the same method fits next. Click any card to open the detail.

Under the hood

How the number is produced.

You never read the whole dataset, and you never take the detector's word for it. Four steps, below. The full walk-through, with the real sample pack, is on the how it works page.

Step 1
Sampleby source

We take a representative sample across the dataset, spread over your sources, so a small read stands in for the whole.

Step 2
Labelyour reviewers

Your reviewers label a small answer key, on your own computers. That is what the number is measured against.

Step 3
Correctfor detector error

We measure how often the detector is wrong, checking it against the answer key, and correct for that, source by source.

Step 4
Reportrate + margin

We report a rate for the whole dataset with a 95% margin of error, and we re-check it as your data changes.

Why a scanner isn't enough

A raw count of matches is not something you can prove.

A scanner flags matches one document at a time. Across a huge dataset that is millions of flags nobody can review, and you cannot stand behind a raw count. We check a sample instead, correct for how often the detector is wrong, and give a rate for the whole dataset with a margin of error. The review cost barely changes as the dataset grows, so the bigger your data, the more this matters. It works the same whether you are counting personal information, copyrighted text, or anything else a person can label.

The overcount

Your redaction number may be too high.

Off-the-shelf detectors are tuned to over-flag rather than miss, so their raw totals overstate how much personal information you actually hold, and you end up redacting, reviewing, and migrating data that was never personal. We correct for that and report only what your reviewers would call personal. On most datasets the honest number lands below the raw scan: less to redact, less to mitigate, a smaller bill. It moves both ways, so when a scanner is under-counting the number goes up instead. Either way, it is the number you can defend.

Scope a diagnostic

Tell us about your data and what you are worried about. We come back with a sampling plan, a labeling budget, and a fixed price.

All demo data is synthetic. We never ship real data out to audit your data.

For personal information the method is tested on public datasets where the true answer was already known (legal judgments and a public privacy dataset), and our number landed on target. It has not yet run on a live customer dataset, and the first project is that check. The other uses above are ones the same method fits but we have not run with you yet. The margin of error covers sampling; it does not settle disagreement over what counts, which we pin down with your reviewers up front.