← Provenote blog
ethics reliability

The asymmetric cost of a false accusation

A missed detection costs you a little certainty. A false positive costs someone else their credibility. Here is the arithmetic almost nobody runs before switching a detector on.

Deepak R Chandran, Ph.D.
Deepak R Chandran, Ph.D.
Founder, Provenote · · 7 min read

Every detector makes two kinds of mistake. It misses AI text that is really there, and it flags human text that was written by a person. Vendors report them together, average them into one number, and call it accuracy. That single number hides the only thing that matters: these two errors do not cost the same, and they are not paid by the same person.

A missed detection costs the institution a little certainty. A false positive costs a specific human being their credibility, in a meeting they did not choose to attend, defending work they actually did.

The arithmetic almost nobody runs

In August 2023, Vanderbilt University published its reasoning for switching its AI detector off. The striking part is not the decision. It is that they simply did the multiplication, in public, using the vendor’s own claimed figure rather than a hostile one.

At a stated false-positive rate of 1%, against the roughly 75,000 papers submitted to the university in 2022, they noted that about 750 papers could have been incorrectly flagged as containing AI-written text. Their conclusion was that they did not believe AI detection software was an effective tool that should be used.

Seven hundred and fifty is not a rounding error. It is a lecture hall full of people, each of whom would have had to prove a negative about their own writing.

And 1% was the vendor’s own advertised figure. Rates vary by tool and by the text they are measured on: in a 2025 University of Chicago Booth working paper, one detector met a stringent policy cap of 0.5% false positives on the authors’ corpus without losing detection performance, while other evaluated tools did not meet the same combined criterion. That is a cap met under their conditions, not a universal error rate. Run the same multiplication against a rate somebody actually measured, on text like yours, and the number moves — potentially a great deal, in either direction. The point is that the institution almost never asks for that rate.

Why a low error rate still produces a coin flip

There is a second effect that a headline accuracy figure conceals completely, and it is the reason a “low” false-positive rate is not reassuring. What matters is not how often the detector is wrong. It is what fraction of its accusations are wrong — and that depends on how much AI text is actually in the pile.

Here is a worked illustration. The inputs are assumptions, chosen to be generous to the detector, not measurements — the point is the shape of the result, not these particular numbers:

A tool with 96% specificity on human text — that is, a 4% false-positive rate — has produced a set of accusations that is close to a coin flip. Note what that figure is not: overall accuracy on this illustration is 95.2%, and neither number tells you what share of the accusations are wrong. Nothing in that calculation is controversial; it is arithmetic. Lower the assumed prevalence of AI text and the picture gets worse, not better, because the innocent pool is larger. This is the number an institution should ask for, and it is almost never the number on the pricing page.

The cost may not be evenly distributed

If the burden fell randomly, it would still be unjust. There is evidence it has not always fallen randomly. In a 2023 study in Patterns, Stanford researchers measured a 61.3% average false-positive rate across seven detectors of that period, on 91 TOEFL essays written by non-native English speakers, while essays by native writers passed. Simplifying the word choices in those essays dropped the rate to 11.6%, which tells you the tools were scoring linguistic sophistication, not authorship. Vanderbilt cited the same concern in its own guidance.

We have to be careful with that finding now. A 2026 study published at the EACL Student Research Workshop retested three detector families and found no systematic bias against non-native speakers, observing that current detectors no longer depend on perplexity in the way the 2023 tools did. It examined Czech rather than English, so it qualifies the earlier result rather than cancelling it. It is a peer-reviewed workshop paper, not an unreviewed preprint. Presenting 2023 as the state of the world in 2026 would be exactly the kind of stale certainty this post is arguing against.

What survives both studies is the part that matters here: whether a given tool is biased against a given population is an empirical question about that tool, on that text, at that moment — and an institution that has not measured it does not know the answer. That uncertainty is itself a reason not to let a score carry a decision.

“We only use it as a signal” does not survive contact with a hearing

The usual defence is that no one treats the score as proof — it is only ever one input among several. That is sincere, and it is not what happens in the room.

A number arrives before the conversation does. “92% AI” anchors everyone present before the student has said a word, and it silently reverses the burden of proof: the person now has to demonstrate they wrote their own essay. Proving that you did write something, after the fact, with no record of having written it, is close to impossible. That is not a process failing occasionally. That is a process working exactly as built, on top of a number that is wrong roughly half the time it accuses someone.

What we do about it

This is the reasoning behind constraints we have deliberately built into Provenote and will not trade away:

There is an honest limit here, and it belongs in the same paragraph as the promise. A verifiable authorship record is strongest when it is created while the work happens. We can examine historical evidence that already exists — repository commits, cloud revision histories, trusted timestamps, signed messages — and sometimes reconstruct part of a document’s history from it. What nobody can do is manufacture an authentic contemporaneous record of events that were never recorded, and we will not imply otherwise — that is precisely the claim this field is already too comfortable making.

The number worth asking for

If you are deciding whether to put a detector between your institution and your students, the question is not “how accurate is it.” Ask two things instead. What is its measured false-positive rate on text like ours? And at our volume, how many innocent people does that produce in a year?

Vanderbilt asked that second question in public and did not like the answer. OpenAI, which trains the models, retired its own text classifier in 2023 for low accuracy — its own reported figures were 26% of AI text caught while 9% of human text was mislabelled. Both are worth more than any marketing figure, because both are organisations publishing something inconvenient about their own product or their own decision.

A missed detection is a missed nudge. A false positive is an accusation. Any system that reports them as one number is hiding the only one that can ruin somebody’s year.
Note

Next up: Disclosure isn’t verification — why saying you used AI and being able to prove how something was made are different things, and why only one of them protects you. See also Principles for the lines we will not cross.

Sources

  1. Vanderbilt University — “Guidance on AI detection and why we’re disabling Turnitin’s AI detector” (16 Aug 2023)
  2. Liang et al., “GPT detectors are biased against non-native English writers,” Patterns (2023)
  3. Al Ali, Helcl & Libovický, “Different Time, Different Language: Revisiting the Bias Against Non-Native Speakers in GPT Detectors,” EACL 2026 Student Research Workshop
  4. OpenAI — “New AI classifier for indicating AI-written text” (with 2023 discontinuation note)
  5. Provenote — Reliability
  6. Provenote — Principles
Deepak R Chandran, Ph.D.
Deepak R Chandran, Ph.D.
Founder, Provenote

Deepak R Chandran, Ph.D., is the founder of Provenote. He writes about content provenance, AI watermarking, and building verification that is honest about what it can and cannot prove.

Get provenance you can prove

Provenote is in private beta. Request early access →

More reading: Absence of a watermark is not proof of a human →