LLMs respond differently to harmful prompts when AI watermarking is used

← Back to the feed

LLMs respond differently to harmful prompts when AI watermarking is used

Ars Technica · 2 weeks ago

Research indicates that SynthID-Text, Google’s AI watermarking system, can alter how language models respond to harmful requests. In some cases, particularly when attackers use prompt-injection techniques, watermarked models may follow instructions they would otherwise refuse, creating potential risks for AI systems and agents.

SynthID-Text subtly changes token selection using a secret key, allowing generated text to be identified as AI-produced without visibly altering it. Researcher Andrea Siposova tested its non-distortionary configuration on six open-weight models and found changes in refusal behaviour, tool selection and potentially the arguments passed to tools, highlighting the need for thorough safety testing.

  • Watermarking can change how models handle harmful prompts.
  • Prompt injection makes the safety risk more pronounced.
  • Developers should test watermarked models and agents thoroughly.

AI Technology

Read the full article at the source →