Skip to content

Anthropic shares more details on how Claude’s new watermarks will work

Anthropo published a blog post Friday seeks to answer some basic questions about how you will watermark text generated by your Claude chatbot. For example: How will the watermark actually work? Can it be hidden with editing? And how does this affect the code?

Claude users have been debating the move since the company revealed earlier this week that would make this watermark to comply with the Transparency Code of the EU AI Law, which requires AI companies to use systems that allow AI-generated content to be identified.

On Reddit, for example, a poster He characterized this as a conspiracy against innocent Claude users.while another claimed: “The only reason you wouldn’t want this is to lie to people.” AND Business Insider reports that “dozens” of X users have claimed to cancel their Claude subscriptions as a result.

Anthropic’s new post begins with an overview of the watermarking concept, explaining that by making “low-risk choices” (such as choosing between the words “cloudy” and “gray” to describe the weather) Claude can create a pattern in his responses that is “undetectable to the reader, but is detectable by anyone who has a key that encodes it.”

“The watermark does not affect the quality of Claude’s production,” the company stated. “To a reader, a watermarked response is indistinguishable from an unwatermarked one.”

More specifically, Anthropic said it will use the SynthID-Text approach that Google DeepMind team described in 2024and that it plans to release a watermark detection API. He also noted that watermarking is distinct from AI detection approaches offered by companies like Pangram who look for “says” in writing (like the construction “what is his is not [X]is [Y]”) to reveal the use of AI: “Detecting these patterns is fundamentally different from looking for a watermark.”

Could someone just rewrite the text to hide the watermark? Anthropic said it’s possible, but “a light edit probably won’t remove the watermark completely,” while “a complete rewrite where every word is replaced will.”

“In the latter case, of course, it is debatable whether the text can still be described as generated by AI,” the company said.

As to whether the watermark will be detectable on text that was only reviewed or edited by Claude, Anthropic said it will depend on “the length of the text and how heavily Claude edited it.” If it has only been lightly edited, “almost every word” will have been written by the human author and “there is very little (if anything) for the watermark to be attached to.”

Meanwhile, the code should be less watermarked than other text, because the model will need to create a working code and will not have the freedom to choose from a variety of equally valid options.

“That said, in areas where there is an arbitrary choice between particular words or terms within the code, watermarking can be used, such as in comments within the code,” Anthropic said. “But by definition it will have a negligible effect on the actual code produced.”

Anthropic also said that Claude will not be the only AI chatbot that will generate watermarked text, as “other major model developers have signed the same Code of Practice and will implement their own watermarks.”

When you purchase through links in our articles, we may earn a small commission. This does not affect our editorial independence.

Leave a Reply

Your email address will not be published. Required fields are marked *