• 4 mins read
  • Published

Developers Race to Strip Claude’s AI Watermarks

Paul Christiano Journalist FAYFO Media

by Paul Christiano

Developers Race to Strip Claude’s AI Watermarks FAYFO Media © fayfo.com
Developers Race to Strip Claude’s AI Watermarks © fayfo.com

Invisible watermarks in AI-generated text are already being removed by coders, just days after new EU rules took effect. Anthropic’s Claude faces rapid circumvention as open-source tools spread. The debate over AI content labeling intensifies.

Just hours after Anthropic confirmed that its Claude AI models would embed invisible, machine-readable watermarks in all generated text, developers began publishing tools to erase them. Guillaume Meyer, a developer, released a workaround within four hours of the announcement. His code, now viral on GitHub, has attracted over 100 contributors and more than 20,000 bookmarks on X, as creators and freelancers rush to integrate the tool into their own projects.

The surge in activity follows Anthropic’s move to comply with the European Union’s AI Act, which requires providers like Anthropic and OpenAI to label synthetic content so it can be detected by machines. The regulation, which took effect earlier this month, threatens fines of up to 3 percent of annual turnover for non-compliance. While the law prohibits model providers from marketing circumvention tools, it does not restrict independent developers from creating or sharing them.

Meyer told WIRED that motivations for bypassing the watermark vary. Some object to mandatory labeling of all AI-generated content, while others, including Meyer himself, are drawn by the technical challenge. He also noted that freelance writers and social media creators have reached out for help using the code. Meyer expressed concerns about the risk of false positives and the inability of watermarks to distinguish between minor and extensive AI assistance, especially for non-native English speakers who rely on tools like Claude and Grammarly for editing.

Anthropic’s watermarking method, based on Google’s SynthID technology, subtly alters word and phrase choices in a way that is invisible to human readers but detectable by machines. Some users worry this could degrade the quality of Claude’s responses, though Anthropic maintains that meaning, quality, and readability remain unchanged. The company plans to release a text-detection API soon, allowing users to check for watermarks themselves.

Meyer’s removal tool works by using a non-watermarking large language model to rewrite the text, swapping synonyms and reorganizing content. This approach depends on using models that do not themselves insert watermarks-a shrinking pool, as over 190 organizations, including OpenAI, Microsoft, and Meta, have signed the EU’s transparency code of practice. Other developers have created similar tools: Erik Hughes built a script in 15 minutes that removes invisible and look-alike characters, reorders sentences, and swaps words for synonyms. Leon Chlon, a Visiting Fellow at the University of Oxford, found that translating Claude’s output into a different dialect and back can also erase the watermark.

Anthropic acknowledges that heavily edited, paraphrased, or translated content may lose its watermark. The company says it is working to improve the system and will soon release its own detection tool, which will allow developers to test the effectiveness of their removal methods. Wayne Pan, cofounder of Haimaker, has already integrated Meyer’s tool into his platform, citing concerns about invisible watermarks being applied even to lightly edited content.

It remains uncertain how many AI labs will implement watermarks on schedule. The EU requires all new models released from August to include watermarks, with existing models to be updated by December. As the debate over AI transparency and content labeling continues, the rapid spread of circumvention tools highlights the technical and regulatory challenges ahead.

Anthropic, founded in 2021 by former OpenAI researchers, has quickly become a major player in the AI space. The company has raised over $7 billion in funding from investors including Google and Amazon, and its Claude models are now used by thousands of organizations worldwide. As regulatory scrutiny intensifies, Anthropic’s approach to compliance and transparency is likely to influence industry standards for AI-generated content.

Related articles