PII detection benchmark

Dated, citable PII-accuracy stats for NeutralAI.

These are the exact numbers NeutralAI publishes for PII detection accuracy: what they measure, how they were produced, and how they compare to a Presidio-vanilla baseline. Vendor-published, not independently audited — methodology is documented below.

Last verified: July 2026

99.8%

public overall F1

NeutralAI measured a 99.8% overall F1 score on its public PII detection benchmark, last verified July 2026.

98.4%

holdout overall F1

On a holdout sample not used for tuning, NeutralAI measured a 98.4% overall F1 score, last verified July 2026.

92.7%

PERSON-entity holdout F1

For the PERSON entity type specifically, NeutralAI measured a 92.7% F1 score on holdout data, last verified July 2026.

57.5%

vanilla Presidio overall F1 (same test set)

On the same benchmark set, an uncalibrated vanilla Microsoft Presidio baseline scored 57.5% overall F1, versus 99.8% for NeutralAI, last verified July 2026.

Methodology

How the numbers were produced.

Dataset

A labeled benchmark set of prompt-style text. The published comparison run covers 1,000 cases across DE, EN, ES, FR, and TR (generated 2026-05-08), scored against supported entity types — names, contacts, financial and account identifiers, and region-specific IDs such as UK NHS numbers.

Holdout discipline

“Holdout” means a sample kept separate from the data used to tune detectors and confidence thresholds. Holdout F1 (98.4%) and PERSON-holdout F1 (92.7%) are reported separately from the public overall F1 (99.8%) so the gap between tuned and unseen performance is visible rather than hidden.

Reproducibility

The methodology is described here: same test set and scorer used for both NeutralAI and the vanilla Presidio baseline, with dataset facts (case count, languages, generation date) published on this page. The labeled dataset and full scoring harness are not public today, so an external team cannot independently rerun this exact comparison. Detailed methodology is available on request through security or partnership review.

This is a vendor-published product benchmark, not an independent or third-party audit. NeutralAI generated, ran, and scored this benchmark itself.See the full Presidio build-vs-buy comparison.

Per-entity results

What granularity actually exists today.

NeutralAI currently publishes overall and PERSON-specific F1 scores. It does not yet publish a full per-entity-type accuracy breakdown (email, phone, card, IBAN, UK NHS number, etc.) — the table below is honest about that gap rather than inventing precision that has not been measured yet.

Entity / scopeMetricScore
Overall (all entity types)Public overall F199.8%
Overall (all entity types)Holdout overall F198.4%
PERSONHoldout F192.7%
UK NHS numberSupported entity type — no published per-entity F1 yet

— = not currently published as a standalone score. See how entity coverage and confidence thresholds work for the supported entity list.

In development — UK legal entity pack

The following UK legal-sector identifiers are on the entity-coverage roadmap. No accuracy numbers exist for them yet — results will be published here once benchmarked.

  • Companies House number
  • HMRC UTR
  • Court references
  • SRA ID
  • Land Registry title number
  • DVLA licence number

UK National Insurance number detection is in development. A gateway recognizer exists, but this entity type is not yet listed in NeutralAI's published supported-entity coverage, so it is not included in the benchmark table above until that coverage is published and measured.

Comparison

NeutralAI vs Presidio vs OpenAI Privacy Filter.

Same benchmark set and scorer for the two measured rows. The OpenAI Privacy Filter row is left explicitly pending — no invented score, no placeholder zero.

Product
Overall F1
Status
NeutralAI
99.8%
Public overall F1 on the NeutralAI PII benchmark set.
Presidio (vanilla, uncalibrated)
57.5%
Same benchmark set and scorer, open-source baseline with no product-layer calibration.
OpenAI Privacy Filter
pending
Not yet evaluated. NeutralAI has not evaluated OpenAI Privacy Filter on this benchmark set — no NeutralAI-verified comparison is published here yet.

FAQ

Honest answers, no spin.

How is it measured?

The headline numbers are F1 scores computed against a labeled benchmark set of prompt-style text, comparing detected PII against ground-truth annotations. The published comparison set covers 1,000 cases across German, English, Spanish, French, and Turkish. Overall F1 covers all supported entity types combined; holdout F1 is measured on a sample not used for tuning the detectors.

Is this independently validated?

No. This is a vendor-published product benchmark, not an independent or third-party audit. NeutralAI generated and scored the benchmark itself. The methodology is documented and the comparison against a vanilla Presidio baseline uses the same test set and scorer so the delta is at least internally consistent.

Can I reproduce it?

Not fully, not yet. The methodology is described on this page, and dataset facts (case count, language split, generation date) are published, but the labeled dataset and full scoring harness are not public — so an external team cannot independently rerun this exact comparison today. Detailed methodology is available on request through a security or partnership review.

How often is this updated?

This page carries a "Last verified" dateline and a changelog. Numbers are refreshed when the underlying gateway benchmark is re-run against a new model or detector release, not on a fixed calendar schedule.

Changelog

Page history.

  • July 2026Initial publication of the standalone benchmark page. UK legal entity pack listed as in development, without figures.
  • May 2026Presidio-vanilla comparison benchmark generated (2026-05-08, 1,000 cases across DE, EN, ES, FR, TR).

Try the masking flow yourself.

These numbers describe detection accuracy. The fastest way to see the product is to run a prompt through it.