How AI guardrails are impeding the work of offensive cybersecurity researchers

← Back to the feed

How AI guardrails are impeding the work of offensive cybersecurity researchers

TechCrunch · 3 hours ago

AI companies’ safety controls, designed to prevent models being used for cyberattacks, are increasingly being criticised by legitimate offensive security researchers and defenders. They argue that the same capabilities needed to test, reproduce and understand vulnerabilities can also be used maliciously, making a clear separation between “offensive” and “defensive” use impractical.

The debate follows US export controls imposed in June on Anthropic’s Mythos and Fable models, reportedly linked partly to concerns that safeguards could be bypassed; Fable 5 returned to general access on 1 July, while Mythos 5 remains available only to vetted US organisations during review. Anthropic and OpenAI run vetting schemes offering some researchers less restricted access, but critics say companies are making arbitrary safety judgements. NCC Group’s Chris Anley said refusals can obstruct vulnerability confirmation, and researchers sometimes turn instead to unguarded open-source models.

  • Cybersecurity safeguards can hinder legitimate vulnerability research.
  • Offensive and defensive AI use is difficult to separate.
  • Researchers may switch to unguarded open-source models.

AI Cybersecurity Research Science Technology

Read the full article at the source →