It’s Frighteningly Easy to Jailbreak Some Frontier AI Models
A report from AI safety nonprofit FAR.AI has found that several leading AI chatbots can be "jailbroken" cheaply and easily, bypassing their safety guardrails to produce harmful content such as cyberattack plans or weapons information. Researchers used an automated tool that generated over a thousand prompt variations to test the defences of major frontier models, raising fresh concerns about the adequacy of AI companies' safety measures and the lack of binding government regulation.
The study tested models from Anthropic, OpenAI, Google and Elon Musk's SpaceXAI. Grok proved most vulnerable, with 448 successful jailbreaks found, followed by Gemini with 249, while Claude, Fable and GPT resisted the automated attacks; jailbreaking Grok cost as little as $58 and Gemini $278. FAR.AI's CEO, Adam Gleave, said AI models are "less regulated than restaurants" and called for external oversight rather than reliance on voluntary industry commitments, though he noted the findings also show systematic safety testing is achievable. Google and Anthropic defended their ongoing safety efforts, while OpenAI and SpaceXAI did not respond to requests for comment; the article also notes patchy US regulation, with only California, New York and soon Illinois imposing safety reporting requirements on AI developers.
- FAR.AI found major AI chatbots can be cheaply jailbroken
- Grok most vulnerable (448 jailbreaks); Claude, GPT, Fable resisted
- Experts urge binding regulation, citing weak US oversight