What Anthropic’s text watermark actually is — and what it can’t tell you
In August 2026 Anthropic began watermarking Claude’s text output. Here is a plain-language read on how it works, what it proves, and where its limits sit.
On August 14, 2026, Anthropic announced that Claude’s text output now carries a statistical watermark in the SynthID-text family, with a detection API described as coming soon. It arrived on 1 September 2026 — in private preview, available to eligible organisations Anthropic lists as regulators, law enforcement, media, fact-checkers, researchers, educational organisations, EU civil society groups, and enterprises with EU AI Act compliance duties. Provenote has not been granted access, and we will say so plainly until we have. It is a meaningful step for provenance — and, like every watermark, it is precise about what it can and cannot establish.
How a text watermark works, briefly
When a language model writes, it repeatedly chooses the next token from a probability distribution. A watermark nudges those choices according to a secret keyed pattern — favouring some tokens over near-equivalent alternatives in a way a reader can’t perceive, but a detector holding the key can measure across a long enough passage.
Detect the pattern and you have a positive, checkable signal that this text came from that watermarking model. That is genuinely valuable: it is a provenance signal you can verify rather than guess at.
What it proves
- A watermark that verifies is strong evidence the text was produced by the watermarking model.
- It is keyed and statistical, so it degrades gracefully — a detector reports a strength, not just yes/no, and needs enough text to be confident.
- It composes with content-credential standards like C2PA, which bind provenance to the file rather than the token stream.
What it can’t tell you
The limits are the important part, because this is where people over-read the result:
- Short text may not carry enough signal to decide either way.
- Paraphrase, translation, and heavy editing can weaken or remove the watermark — deliberately or just as a side effect of normal revision.
- Other models (open-weight, older, or simply un-watermarked) leave nothing to find.
- Therefore no watermark found ≠ written by a human. It is the expected result for most text on the internet.
A watermark is a strong yes and a weak no. Treat the yes as evidence; never treat the no as a verdict.
Where Provenote sits
Provenote is the neutral verification layer over these signals. When Anthropic’s detection API ships, we plug it in as one trusted provenance source alongside SynthID and C2PA — reading the positive signal when it’s there, and being honest about uncertainty when it isn’t. We are watching for that launch and will light it up as soon as it’s available.
Three limits Anthropic states itself, which matter more than the announcement. The watermark is going into future Claude models, with older models being covered over the following months — so its absence today may simply mean the model predates the rollout. It says nothing about ownership or authorship. And it cannot distinguish “Claude wrote this” from “Claude heavily edited this.” A provider being that precise about what its own signal does not prove is the standard the rest of this field should be held to.
The EU AI Act’s Article 50 transparency duties came into force on 2 August 2026, which is exactly why provider watermarks are arriving now. Generative systems already on the market get until December 2026 to add machine-readable marking; anything placed on the market since August has had to mark from day one. Reading those marks correctly — and not over-claiming — is the whole job.
Want the mechanics of how we weight each signal into a single calibrated verdict? That’s our Methodology, and the next post below goes deeper on measuring detector error honestly.
Sources
Deepak R Chandran, Ph.D., is the founder of Provenote. He writes about content provenance, AI watermarking, and building verification that is honest about what it can and cannot prove.
Provenote is in private beta. Request early access →