A recent study by Lasso has shown that digital watermarks used in AI systems, such as Google DeepMind’s SynthID-Text, might unintentionally change the behavior of large language models (LLMs). Researchers tested how these watermarks affect AI behavior and found that they can influence how models respond to harmful requests, their susceptibility to prompt injection attacks, and the tools they choose to use. One major concern is that even without an intentional attack, the watermark may cause models to be more likely to comply with potentially harmful instructions, altering their normal behavior.
AI companies are increasingly using invisible digital watermarks to label content created by their systems. This is partly to meet new regulations, such as the European Union's AI Act, which requires greater transparency in AI-generated content. Google’s SynthID watermark is known for being hard to alter or compress, but experts say it is not a perfect solution for preventing misinformation. Challenges include the lack of compatibility between different platforms and the existence of open-source AI models that lack such protections.
Anthropic, the company behind the AI assistant Claude, has announced plans to use a similar watermarking system in future versions of its product. The company also intends to embed digital provenance metadata into image files, following the C2PA standard. This metadata would help identify AI-generated content and is part of its efforts to comply with the EU’s transparency rules. However, experts warn that these measures are not foolproof—metadata can be stripped away, and watermarks can become harder to detect.
Lasso’s research highlights a key issue: AI watermarks may not only be ineffective in stopping misuse but could also change how AI systems behave. The study found that the watermarking process might cause models to produce different outputs or make different decisions, a phenomenon called "sampling drift." This behavior was observed even without an attack, and when combined with prompt injection—where users trick AI into following harmful instructions—the effects were more severe. Despite these risks, companies like Anthropic continue to adopt watermarking technology, emphasizing its role in traceability and compliance.
In response to these developments, a group of developers recently released a tool called "Remove-AI-Watermarks," designed to strip AI-generated content of both visible and invisible identification markers. The tool uses advanced image processing techniques to bypass automatic detection systems on social media. The developers note that their tool is intended for research and privacy protection, not for malicious use. However, they also caution about the legal risks associated with using such technology to deceive or mislead.
AI Digital Watermarks May Influence Model Behavior, Study Finds
AI-rewritten from original reportingHow it works
ai-watermarksllm-behaviorsynthid-textanthropiceu-ai-actsecurity-concerns



