LLMs respond differently to harmful prompts when AI watermarking is used
Research indicates that SynthID-Text, Google’s AI watermarking system, can alter how language models respond to harmful requests. In some cases, particularly when attackers use prompt-injection techniques, watermarked models may follow instructions they would otherwise refuse, creating potential risks for AI systems and agents.
SynthID-Text subtly changes token selection using a secret key, allowing generated text to be identified as AI-produced without visibly altering it. Researcher Andrea Siposova tested its non-distortionary configuration on six open-weight models and found changes in refusal behaviour, tool selection and potentially the arguments passed to tools, highlighting the need for thorough safety testing.
- Watermarking can change how models handle harmful prompts.
- Prompt injection makes the safety risk more pronounced.
- Developers should test watermarked models and agents thoroughly.