How Claude's Invisible Watermark Works

On August 2, 2026, Anthropic began embedding invisible watermarks in all text generated by Claude. Here's a technical breakdown of how the system works.

Token-Level Statistical Watermarking

Claude's watermark operates at the token selection level during text generation. When Claude generates text, it doesn't just pick the most likely next word — it selects from a probability distribution of possible tokens. The watermarking system subtly biases these selections in a pattern that's invisible to humans but statistically detectable.

Think of it like a card dealer who appears to shuffle fairly but actually arranges cards in a pattern only they know. Each individual word looks natural, but across hundreds of tokens, the pattern emerges.

How the Pattern Survives

Unlike image watermarks that can be removed with a screenshot, the text watermark survives:

What Breaks the Watermark

The watermark IS vulnerable to:

C2PA Content Credentials

In addition to the text watermark, Claude signs generated files (images, PDFs) with C2PA metadata. C2PA (Coalition for Content Provenance and Authenticity) is an open standard that embeds provenance information in file headers.

Unlike the text watermark, C2PA metadata is trivially removable — converting the file format, re-saving, or taking a screenshot strips the header entirely.

Why Anthropic Added Watermarks

The EU AI Act requires AI-generated content to be identifiable. Anthropic's watermarking system was introduced to comply with these regulations, initially for EU users but now applied globally to all Claude output.

Anthropic has stated the watermark is designed to be "robust but not unbreakable" — acknowledging that determined users can remove it while making casual AI-generated content detectable.

Remove Watermark → Check Your Text →