Detector performance varies enormously between tools. Some come close to their claims. Many do not. So the only number worth anything is the one measured on text like yours. We publish what we measure, including our own limits, and what we refuse to claim.
Because we hold the keys, every sample has a known label. On a small labeled corpus generated with our own SynthID watermark, our detector recovered the truth cleanly. This validates the pipeline, it is not a claim about any provider’s production watermark.
Method: gpt-2 generations, SynthID-Text watermark, mean-g-value detector, own keys. Corpus scales as the harness grows.
| Measure | Marketed | Independently measured | What it means |
|---|---|---|---|
| Overall accuracy | 99%+ | varies sharply by tool | In a 2025 University of Chicago Booth study, one detector reached 99.8, 100% accuracy while other commercial tools performed materially worse. Accuracy is a property of the specific tool, not of the category. |
| False-positive rate | “under 1%” | ≤0.5% for the best tool tested | That same study found only one tool holding false positives at or below 0.5% without sacrificing accuracy. Others did not. Ask any vendor for the rate measured on text like yours. |
| Non-native English writers | , | contested | A 2023 Patterns study found >50% of TOEFL essays by non-native writers wrongly flagged by the detectors of that period. A Feb 2026 preprint retesting three detector families (in Czech) found no systematic bias. Treat bias as a per-tool, per-population question to be measured, rather than a settled fact either way. |
In Provenote, every statistical detector is a Tier-C input weighted by its measured FPR, the higher that rate, the smaller its say, and it is a lead either way, never a verdict. See it in a sample report →
Sources: Jabarian & Imas, “Artificial Writing and Automated Detection,” University of Chicago Booth (2025), SSRN; Liang et al., Patterns (2023), paper; Al Ali, Helcl & Libovický, preprint (Feb 2026), arXiv; Vanderbilt University guidance (2023), statement; OpenAI classifier note (2023), announcement. Every figure on this site names the study it came from, the population it was measured on, and its date. Where the evidence conflicts we say so, rather than picking the flattering number. We do not cite vendor-published comparisons as independent evidence. We deliberately do not name individual detector vendors: our argument is about what the category delivers, not about any one company. We do name the standards and watermarks we read, SynthID, C2PA, Anthropic’s Claude watermark, because you should know exactly which signals a verdict rests on. Figures are indicative and dataset-dependent, that is precisely the point.