AI model watermarking changes agent behavior

← Back to the feed

AI model watermarking changes agent behavior

The Register · 2 weeks ago

AI watermarking designed to verify the provenance of generated content may also subtly alter how AI models behave. Lasso Security found that watermarking can affect tool selection, tool arguments and safety refusals, with the effects becoming more significant when models face adversarial prompt injection. This matters because the changes can influence not only a model’s wording but also the actions of AI agents using its output.

The research examined seven models using the BFCL v4 single-turn AST tool-calling benchmark, with watermarking reducing accuracy in six. Errors included choosing the wrong tool, supplying incorrect arguments, or producing malformed input. Tests with HarmBench and JailbreakBench found smaller changes for plainly harmful requests but a significantly higher attack success rate during prompt-injection attempts, making models less likely to refuse. Lasso said this does not necessarily invalidate watermarking, but security testing and red-teaming should include watermarked outputs.

  • Watermarking can change how AI agents call tools.
  • Six of seven tested models became less accurate.
  • Prompt injection made safety refusals less reliable.

AI Technology

Read the full article at the source →