← Provenote blog
standards provenance

Content Credentials for text: what C2PA does — and doesn’t — cover

C2PA is the “nutrition label” for digital media, and it now reaches into documents. But there’s one thing it fundamentally can’t do for writing — and knowing that line is the whole point.

Deepak R Chandran, Ph.D.
Deepak R Chandran, Ph.D.
Founder, Provenote · · 6 min read

C2PA — “Content Credentials” — is the closest thing we have to a nutrition label for digital media: a signed record of how a file was made and edited. It now reaches into documents, too. But there is one thing it fundamentally cannot do for writing, and knowing that line is the whole point.

What C2PA is

A cryptographically signed manifest bound to a file, recording its origin and edit history — tamper-evident provenance that anyone can verify against a published trust list. It began with images and spread to audio and video.

It reaches text now

Text support arrived in stages, and the stages matter. 2.3 (December 2025) added embedding into unstructured text files. 2.4 (April 2026) added HTML documents and structured text formats — source code, YAML, Markdown, AsciiDoc — alongside the document containers (PDF, Office/OOXML, ODF, EPUB and other ZIP-based formats) the spec already covered. A document can now carry a verifiable Content Credential.

2.4 also introduced a dedicated AI Disclosure Assertion (c2pa.ai-disclosure) — machine-readable AI transparency information carried inside the credential itself. That is a standards body building the disclosure signal directly into the provenance record, which is precisely the kind of authoritative evidence we would rather read than infer.

The catch — and it is the important part

For a credentialed PDF, the manifest binds to the file rather than to the words. Copy the sentences out and paste them into an email, and the credential does not come with them: the provenance was attached to the container, not the prose.

That is not the whole story, and the correction matters. Specification 2.4 Annex A.8 defines embedding a manifest into unstructured text itself, using Unicode variation selectors interleaved with the characters, with A.9 doing the same for structured text. A credential built that way lives in the text, not in a container around it. The specification scopes the method narrowly, and the scope is the point: it is for “unstructured text where traditional file-based embedding is not practical”, and it offers “content intended for copy-paste operations across different systems” as its example of that, the aim being to ensure “Content Credentials persist with the content itself across platforms”.

Three honest qualifiers belong with that, and two of them are the specification’s own, in the same paragraph. It says the approach “should only be used with unstructured text assets where no other embedding method is feasible” — a last resort, not a preferred route. It says the approach “remains under review and may be subject to change based on implementation feedback and interoperability testing”. Ours is the third: we have not tested whether such a credential survives any particular copy, paste or normalisation, so we do not claim that it does. What the standard intends and what a given editor preserves are different questions.

What the standard defines, and what ships

Everything above is what the specification defines. It is not what you can buy. We measured the C2PA Conforming Products List on 24 September 2026: of 219 conformant products, not one declares a non-empty text media type under the three text keys the schema defines. The Conformance Program had published a text asset conformance rubric seven weeks earlier, on 6 August 2026, and the public test corpus contains no text assets at all. The grading instrument exists, the assets to grade against it do not, and no product has yet presented itself for grading.

We report that as a dated census rather than a verdict: the list changes, and we would be glad to be out of date. The full measurement, with the script that produces every number in it and the content hashes of every source it read, is published at doi:10.5281/zenodo.22969642.

A watermark travels with the words. A content credential usually travels with the file, and only as far as the next tool that knows to carry it. The standard now defines a third option, a credential embedded in the characters themselves, which is newer than most software that will handle it.

Why you want more than one signal

A text watermark rides along with the tokens: it survives copy-paste but is fragile to heavy rewriting. A C2PA credential on a document rides with the file: robust while the file is intact, and not carried by the text once it is extracted. Annex A.8 offers a third shape, a manifest embedded in the characters, which is why the neat opposition between the two is less neat than it looks. Because they break in different ways, a provenance-first verifier reads whatever is there. Provenote validates real C2PA manifests as a first-class signal.

Note

Next up: the asymmetric cost of a false accusation — why we cap what a statistical detector is ever allowed to conclude. See the signal stack →

Sources

  1. C2PA — Content Credentials
  2. C2PA Technical Specification 2.4 (April 2026) — HTML + structured text embedding, AI Disclosure Assertion
  3. Conforming Products List — the census in this post, measured 2026-09-24
  4. C2PA Text Asset Conformance Rubric, v1.0, 6 August 2026
  5. Chandran, D. R. (2026) Text Provenance in Practice — the full census, with the script that produces every number in it
  6. Provenote — Methodology
Deepak R Chandran, Ph.D.
Deepak R Chandran, Ph.D.
Founder, Provenote

Deepak R Chandran, Ph.D., is the founder of Provenote. He writes about content provenance, AI watermarking, and building verification that is honest about what it can and cannot prove.

Get provenance you can prove

Provenote is in private beta. Request early access →

More reading: Absence of a watermark is not proof of a human →