Reliability

Reliability, in the open.

Detector performance varies enormously between tools. Some come close to their claims. Many do not. So the only number worth anything is the one measured on text like yours. We publish what we measure, including our own limits, and what we refuse to claim.

Our own detector · measured

SynthID watermark: generate → detect, on our own keys

Because we hold the keys, every sample has a known label. On a small labeled corpus generated with our own SynthID watermark, our detector recovered the truth cleanly. This validates the pipeline, it is not a claim about any provider’s production watermark.

1.00
True-positive rate
0.00
False-positive rate
0.52
Decision threshold (mean g-value)
n=12
Labeled samples, early, small

Method: gpt-2 generations, SynthID-Text watermark, mean-g-value detector, own keys. Corpus scales as the harness grows.

Independent evidence · statistical detectors

What statistical detectors actually deliver, marketed vs. measured

MeasureMarketedIndependently measuredWhat it means
Overall accuracy 99%+ varies sharply by tool In a 2025 University of Chicago Booth study, one detector reached 99.8, 100% accuracy while other commercial tools performed materially worse. Accuracy is a property of the specific tool, not of the category.
False-positive rate “under 1%” ≤0.5% for the best tool tested That same study found only one tool holding false positives at or below 0.5% without sacrificing accuracy. Others did not. Ask any vendor for the rate measured on text like yours.
Non-native English writers , contested A 2023 Patterns study found >50% of TOEFL essays by non-native writers wrongly flagged by the detectors of that period. A Feb 2026 preprint retesting three detector families (in Czech) found no systematic bias. Treat bias as a per-tool, per-population question to be measured, rather than a settled fact either way.

In Provenote, every statistical detector is a Tier-C input weighted by its measured FPR, the higher that rate, the smaller its say, and it is a lead either way, never a verdict. See it in a sample report →

Sources: Jabarian & Imas, “Artificial Writing and Automated Detection,” University of Chicago Booth (2025), SSRN; Liang et al., Patterns (2023), paper; Al Ali, Helcl & Libovický, preprint (Feb 2026), arXiv; Vanderbilt University guidance (2023), statement; OpenAI classifier note (2023), announcement. Every figure on this site names the study it came from, the population it was measured on, and its date. Where the evidence conflicts we say so, rather than picking the flattering number. We do not cite vendor-published comparisons as independent evidence. We deliberately do not name individual detector vendors: our argument is about what the category delivers, not about any one company. We do name the standards and watermarks we read, SynthID, C2PA, Anthropic’s Claude watermark, because you should know exactly which signals a verdict rests on. Figures are indicative and dataset-dependent, that is precisely the point.

What we do not claim

  • • We cannot prove a text is human-written. No tool can.
  • • Absence of a watermark is not evidence of anything.
  • • Statistical detectors are corroboration, never proof, and never grounds for an accusation on their own.
  • • Our own harness is early and small-N; the numbers grow and change as the corpus does, and we publish them as they do.
  • • Provider production watermarks require the provider’s own detection API; Claude’s is live but in private preview, and we have not been granted access.