skip to content
The Weighted Average

AI Safety & Security

Claude Watermarks Make Provenance a Runtime Duty

Anthropic plans invisible text watermarks and C2PA file metadata; teams must preserve detection and disclosure paths.

A printed document resting on a wooden desk
A printed document resting on a wooden desk. Photograph by Lewis Keegan

Anthropic says Claude will mark generated text with an imperceptible watermark and attach signed C2PA provenance metadata to supported files. The company is introducing two complementary techniques, so teams shipping Claude output should treat provenance as a runtime obligation: preserve the mark where possible, disclose the model path, and keep a fallback when detection cannot establish origin. The AMIE Video lead makes the parallel clinical point: a system’s evidence and escalation path are part of the product, not an afterthought.

Invisible is not the same as optional

Anthropic’s support guidance on Claude-generated content says it has signed the EU AI Act’s Article 50(2) Code of Practice on transparency of AI-generated content. The page describes a plan rather than a fully specified production protocol. Text from supported Claude models will carry an imperceptible watermark intended to travel with copied text and persist through some editing. Supported generated files will carry signed provenance metadata using the C2PA open standard. The Verge report on the rollout adds that Anthropic describes the change as a future commitment and is still working through existing-model support.

The two channels behave differently. Text moves into email, source files, tickets, documents, and code review; a model-level mark is intended to follow it across Claude surfaces. A file carries metadata that can be signed and checked for tampering, but a conversion or upload can remove it.

Anthropic names three supported file types in its explanation—.svg, .png, and .jpg—while saying “such as,” so that list is illustrative rather than a complete coverage matrix. The product decision is not to assume every MIME type is marked. A team needs a per-surface inventory: model and version, output format, export path, transformation step, storage system, and the point at which a detector or disclosure label is applied.

A detected mark indicates content may have been processed by Claude, not that every sentence was written by Claude. Content can be edited, combined, translated, or passed through several systems. Provenance signals processing history; absence of a mark does not prove human origin.

The regulatory direction is clear enough to change planning. The European Commission’s AI Act framework describes transparency obligations for certain AI-generated or manipulated content, including machine-readable marking and disclosure in relevant cases. Anthropic’s support page says its implementation is designed to support Article 50 commitments, but also tells developers to assess what the rules require of their own products and services. A vendor mark does not transfer the deployer’s legal analysis to the model provider.

That boundary should guide procurement. Ask whether coverage includes text, images, code, audio, and transformed documents; whether marks survive API, partner-cloud, and IDE paths; and what happens when downstream systems strip metadata. “The provider marks output” is a capability statement. “Our service can make a defensible disclosure decision” is a control.

The earlier Claude watermark Wire note caught the announcement. Provenance has moved from a policy page into the application boundary; teams need not reject Claude, but they must own the output path.

Google’s C2PA content-credentials work shows that provenance is becoming a cross-vendor interface rather than an Anthropic-only promise. That makes the file standard useful, but it also raises the bar for interoperability: a content-management system should preserve a manifest whether the asset came from Claude, Gemini, or another supported generator.

Provenance fails at the handoff

C2PA gives file provenance a shared vocabulary. The C2PA specification project describes an open technical standard for recording the history and provenance of digital content. Anthropic says its signed metadata can indicate that a file was processed by Claude and can help detect whether the file was tampered with. That is valuable at an upload boundary, an editorial review queue, or an asset-management system that can preserve the manifest.

The hard part is the chain after generation. A design tool may flatten an SVG into a raster image. A content-management system may recompress a JPEG. A user may copy the text into a plain-text field. A document converter may preserve the visible content while dropping metadata. The mark can be technically sound at the origin and operationally absent at publication. The right metric is therefore not “does Claude watermark?” but “what percentage of outputs reach the decision point with a recoverable provenance signal?”

Teams should instrument that path. Capture the source model, surface, timestamp, job identifier, and transformation steps in an audit record. Preserve C2PA manifests when a file supports them. For text, store a disclosure event and detector result rather than trying to inspect invisible marks inside every downstream string.

Anthropic says it is still working on detection mechanisms and will share technical documentation later. Until a detector is available and tested, a product cannot promise that every generated item can be verified after editing. It can still disclose Claude assistance, restrict high-risk uses, and retain the generation event. The gap calls for conservative process design, not denial.

A two-technique design creates a useful planning figure: one model-level marking layer spans the Claude surfaces Anthropic names, while the second attaches to files. The distinction is operationally important. One control follows model output across surfaces; the other depends on file format and every handoff that preserves it.

The strongest counterpoint is that watermarking may be too fragile to justify much engineering. Anthropic does not yet publish detection robustness, false-positive rates, persistence after editing, or the exact text method. C2PA metadata can be stripped, and a watermark that survives only friendly transformations could become a checkbox rather than evidence.

That skepticism is warranted, but it does not restore the old default. Teams already make authorship and disclosure decisions with weaker evidence. A provenance signal attached at the model boundary can improve the audit even if it is not infallible. The comparison is not perfect detection versus no detection; it is a measured signal plus a fallback versus an undocumented generation path.

The decision should change when Anthropic publishes a detector, persistence tests across common transformations, supported-model coverage, and independent error evaluations. It should change in the other direction if routine editing destroys the text signal, metadata disappears in production tools, or users treat “no mark” as proof of human authorship. Until then, keep provenance as a positive signal and never as a negative certificate.

For the next quarter:

  • Product teams shipping Claude output should create a provenance inventory before enabling a new surface. Record which technique applies, which formats are supported, and where metadata can be lost.
  • Platform teams should preserve generation events outside the artifact. Store model, version, surface, timestamp, and transformation metadata under a retention policy that matches the content’s sensitivity.
  • Compliance teams should prohibit “no mark means human” decisions. Require an explicit disclosure or human-review path for content where origin matters; Anthropic says detection is forthcoming.
  • Procurement should demand a detector and test corpus. Run copied text, edited text, screenshots, format conversions, and API or partner-cloud paths through the proposed workflow.

The Pangram detection-cost analysis showed why a detector’s score and its price are not enough; buyers need a corpus-specific error budget and an appeals path. Anthropic is now offering a different kind of signal. It will be useful only if the application around it preserves the chain of custody.

Sources