Claude’s New Watermarks Explained by Anthropic

Understanding the Concept of Watermarking in AI-Generated Text
Anthropic recently published a blog post detailing its approach to watermarking text generated by its chatbot, Claude. The company aims to address key questions about how this process works, whether it can be hidden through editing, and what impact it has on code generation.
The move comes as part of Anthropic’s compliance with the EU AI Act’s Transparency Code, which requires AI companies to implement systems that make it possible to identify AI-generated content. Since the announcement, there have been discussions and debates among users, particularly on platforms like Reddit and X, regarding the implications of this change.
On Reddit, some users have expressed concerns, with one poster characterizing the initiative as a “conspiracy against innocent Claude users.” Others have suggested that the only reason someone would oppose this is to “lie to people.” Meanwhile, Business Insider reported that “dozens” of users on X have claimed to cancel their Claude subscriptions due to this change.
How Does the Watermarking Work?
In its blog post, Anthropic begins by explaining the concept of watermarking. It states that when making “low-stakes choices,” such as selecting between the words “overcast” and “grey” to describe the weather, Claude can create a pattern in its responses that is “undetectable to the reader, but is detectable to anyone who has a key that encodes it.”
The company emphasizes that watermarking does not affect the quality of Claude’s output. To a reader, a watermarked response is indistinguishable from an unwatermarked one.
Technology Behind the Watermark
More specifically, Anthropic revealed that it will be using the SynthID-Text approach, which was outlined by the Google DeepMind team in 2024. The company also plans to release a watermark detection API. It further clarified that watermarking is different from other AI detection methods used by companies like Pangram, which look for specific patterns in writing (such as the construction “his isn’t [X], it’s [Y]”) to identify AI usage.
“Picking up on these patterns is fundamentally different from checking for a watermark,” the company stated.
Can the Watermark Be Hidden?
One of the key concerns raised by users is whether the watermark can be hidden through editing. Anthropic acknowledged that it is technically possible to rewrite the text to hide the watermark, but “light editing probably won’t remove the watermark completely.” However, a complete rewrite where every word is replaced would likely eliminate the watermark.
“In the latter case, of course, it’s arguable whether the text can any longer be described as AI-generated,” the company said.
Impact on Proofread or Edited Text
Regarding whether the watermark will be detectable in text that has been proofread or edited by Claude, Anthropic explained that this depends on “the length of the text and how heavily Claude has edited it.” If it’s only been lightly edited, “nearly all the words” will have been written by the human author, and “there’s very little (if anything) for the watermark to attach to.”
Watermarking in Code Generation
When it comes to code generation, the watermark should have less of an effect than in other types of text. This is because the model must produce working code and won’t have the freedom to choose between various equally valid options.
“Having said that, in areas where there is an arbitrary choice between particular words or terms within the code, the watermark can be used, such as comments within code,” Anthropic noted. “But by definition, it will have a negligible effect on the actual code produced.”
Broader Implications
Anthropic also mentioned that Claude won’t be the only AI chatbot to generate watermarked text. The company stated that “other major model developers have signed the same Code of Practice and will be implementing their own watermarks.” This suggests that watermarking is becoming a standard practice across the AI industry.























