Research

How Hexagone AI performs.

Every anonymization vendor claims high accuracy. Very few say on which dataset, against which metric, or compared to what. This page collects the evaluations we have run and published, with the methodology behind each and the papers and reports to read in full. If you find a mistake, tell us and we will correct the page.

Headline results

97.9%
of annotated personal values removed on PII-INDEXBENCH
Measured over the whole dataset rather than a sample of it: 12,079 documents and 226,428 annotated values. The next system, an open-weights Qwen3-14B prompted as a detector, found 91.7%, and precision on the same run was 93.5%.

Measured on: PII-INDEXBENCH v1.2, every split · September 2026

95.3%
of annotated personal values removed, across six languages
47,566 documents in English, French, German, Spanish, Italian and Dutch, and the best recall of the five systems compared. The spread between the strongest and weakest language is 1.3 points.

Measured on: AI4Privacy PII-Masking-300k, full validation split · September 2026

14%
re-identification rate at level 1, the lowest of the tools compared
Azure scores 22% and a GPT-4.1 purifier prompt 26% on the same texts. At level 2, where identifiers are obfuscated: 27% against 33% and 58%.

Measured on: RAT-Bench test set · Imperial College London

0
direct identifiers recovered from the anonymized texts
No name, phone number or social security number was recoverable at either difficulty level. Azure and the GPT-4.1 prompt both leaked some.

Measured on: RAT-Bench test set · Imperial College London

These figures describe performance on the public test sets named above, and they are not a guarantee about your own documents. A different corpus gives different numbers, which is why every result here names the corpus it was measured on.

The studies

  1. Full-dataset evaluation

    PII-INDEXBENCH: 12,079 documents built to be hard

    Every record through the production API, scored against the benchmark's own annotations alongside five other systems, one of which is Qwen3-14B prompted as a detector. The precision figure is taken apart against those same annotations further down the page.

    Run by Hexagone AI, September 2026. Dataset PII-INDEXBENCH v1.2, CC BY 4.0.

    97.9%
    recall, highest of six systems
    93.5%
    precision
    See the results
  2. Full-dataset evaluation

    AI4Privacy: 47,566 documents in six languages

    The same protocol as PII-INDEXBENCH, applied to the most cited public PII corpus and scored per language and per kind of personal data. It also comes with the caveats that follow from a corpus this widely distributed.

    Run by Hexagone AI, September 2026. Public dataset, reproducible.

    95.3%
    recall, highest of five systems
    6
    languages, 1.3 points apart
    See the results
  3. Independent benchmark

    RAT-Bench: can an attacker still find the person?

    A re-identification benchmark built at Imperial College London, and a better question than most evaluations ask. We ran our pipeline against it alongside Azure and a GPT-4.1 masking prompt, using the authors' own scoring criteria unchanged.

    Benchmark by Krčo, Yao, Meeus and de Montjoye, Imperial College London, 2026

    14%
    re-identified at level 1, the lowest compared
    0
    direct identifiers recovered
    0.83
    BLEU: how much text survives
    See the results
  4. In-house evaluation, 2024

    Detection rate against Microsoft Presidio

    Our first published comparison, from November 2024, and the one that comes with a ten-page methodology PDF. Superseded in scope by the two full-dataset runs above, and kept here because its numbers still stand.

    Run by Hexagone AI, November 2024. Public dataset, reproducible.

    96.95%
    of entities detected, against 75.30%
    93.34%
    on degraded text, against 65.89%
    See the results

How to read any vendor benchmark. Ask three questions. Which dataset, and is it public? Which metric, and does it measure risk or just recall? Which systems were compared, and were they configured fairly? A number without those three answers, ours included, is marketing. Every study above carries all three.

Run it on your own documents.

A benchmark is someone else's data. One week free, no credit card, and the files never leave your machine, so there is nothing to risk in checking for yourself.