The watermark lives in word choice, not hidden characters
Language models repeatedly choose among plausible next words. The watermark uses a secret key and the preceding context to influence the random choice among options that should preserve meaning and quality.
A reader sees ordinary prose. A detector with the key can examine a longer sequence and estimate whether its pattern is consistent with Claude's watermarked generation. Nothing is inserted as a hidden character, and no user identity is encoded.
The detector answers a narrow question
Anthropic says detection can estimate whether Claude was partly involved. It cannot confirm that a human wrote unmarked text, identify a different AI system, or tell whether Claude generated a passage from scratch or heavily edited it.
The method also has less evidence in short, factual, or lightly modified passages. Exact code and constrained facts leave fewer harmless word choices in which a watermark can operate.
Do not turn probability into punishment
Schools, employers, publishers, and courts may be tempted to treat a positive detector result as proof of misconduct. Anthropic's own limits argue against that use. A score must be accompanied by sample length, confidence, editing context, and an opportunity to challenge the conclusion.
Watermarking can improve provenance across a crowded information system. Its credibility depends on resisting claims the signal was never designed to prove.
- Publish confidence, sample length, and known failure modes with every result.
- Do not infer human authorship from the absence of Claude's watermark.
- Require corroborating evidence before disciplinary or legal action.
- Keep text watermarks distinct from C2PA credentials attached to supported files.
Go to the source
Read the evidence behind this analysis. External links open in a new tab.
Anthropic — How Claude's text watermark works Nature — Scalable watermarking for identifying large language model outputs European Commission — Transparency code for AI-generated content


