Grading detectors honestly: how we measure false-positive rates
A detector’s advertised accuracy is the wrong number. Here’s the number that actually matters, what it costs a real person when it’s wrong, and how Provenote grades every detector it touches.
Every AI detector leads with one figure: “99% accurate.” It is the number on the pricing page, in the sales deck, and the one a nervous administrator repeats in a misconduct hearing. On its own it is close to meaningless — and building a decision on it is how honest people get hurt.
Accuracy hides the harm
“Accuracy” blends two very different things: how often a detector catches AI text, and how often it wrongly flags human text. Those failures are not equal. Catching a little less AI is a missed nudge. Flagging a human is an accusation. The number that lands on a real person is the false-positive rate — the share of genuinely human writing a detector calls machine-made — and it is usually buried, because it is the number that looks bad.
OpenAI knows this better than anyone. It built its own AI Text Classifier in early 2023 and retired it about six months later “due to its low rate of accuracy.” Its own reported figures: it correctly flagged just 26% of AI-written text — while mislabelling 9% of human text as AI. The company that trains the models could not ship a text detector it trusted. That should end the “99%” conversation.
The cost is asymmetric — and unequal
A false positive is not a rounding error in a benchmark. It is a student in a disciplinary meeting, a job applicant whose essay was auto-rejected, a writer told to prove they wrote their own words.
The harm may also not be evenly spread. In a 2023 study in Patterns, Stanford researchers ran the detectors of the day over essays by non-native English writers and found more than half of the TOEFL essays were misclassified as AI-generated, while native-writer essays sailed through. The proposed mechanism was mundane: simpler, less “surprising” prose read as machine-made to tools that scored text on how predictable it was.
That evidence is now contested, and we are not going to present a 2023 result as today’s reality. A February 2026 preprint retested three detector families and reported no systematic bias against non-native speakers, noting that contemporary detectors no longer lean on perplexity the way earlier ones did. It tested Czech rather than English, and it has not yet been peer-reviewed, so it qualifies the 2023 finding rather than overturning it.
The honest position is therefore narrower and more useful than either headline. Bias is a property of a specific tool on a specific population at a specific time — not a permanent law of detectors, and not a solved problem either. That is precisely why we insist on a measured false-positive rate for the text you actually care about, rather than any general reassurance.
A missed detection costs you a little certainty. A false positive costs someone else their credibility. Those are not the same mistake, and they should never share a headline number.
How Provenote grades a detector
We treat every statistical detector as a witness to be cross-examined, not an oracle:
- We measure each detector’s false-positive rate on held-out human text we control — not the vendor’s self-reported figure.
- We weight its vote by that measured rate, not by its marketing. The higher a detector’s measured false-positive rate, the smaller and more clearly-bounded its say — and it is never a verdict at any rate.
- Every report shows the gap between three numbers: the detector’s raw score, its measured false-positive rate, and the trusted weight we actually give it. That gap — between what a tool claims and what it earns — is the whole product.
- A statistical detector, on its own, can never move a Provenote result to “confirmed.” Confirmation comes from provenance — a verifiable watermark or credential — or it does not come at all.
What good looks like
Provenance first; detectors corroborate. A detector can add weight to a picture that verifiable signals already support. It cannot, by itself, convict — because we have measured exactly how often it would convict the innocent. Honest grading is less flattering than “99% accurate.” It is also the only version that is safe to put in front of a decision that affects a real person.
Next up: Article 50, in plain English — what the EU AI Act’s transparency duties, in force since 2 August 2026, actually require, and what the December 2026 deadline really covers. See how we measure detectors →
Sources
- OpenAI — “New AI classifier for indicating AI-written text” (with 2023 discontinuation note)
- Liang et al., “GPT detectors are biased against non-native English writers,” Patterns (2023)
- Al Ali, Helcl & Libovický, “Different Time, Different Language: Revisiting the Bias Against Non-Native Speakers in GPT Detectors,” preprint (Feb 2026)
- Provenote — Reliability
- EU AI Act — Article 50
Deepak R Chandran, Ph.D., is the founder of Provenote. He writes about content provenance, AI watermarking, and building verification that is honest about what it can and cannot prove.
Provenote is in private beta. Request early access →