OpenAI Is About to Release Its First AI Model With ‘Critical’ Cyber Abilities

← Back to the feed

OpenAI Is About to Release Its First AI Model With ‘Critical’ Cyber Abilities

Wired · 2 hours ago

OpenAI says its forthcoming model, Astra, has crossed the "critical" cybersecurity capability threshold set out in its own preparedness framework, meaning it can independently find and exploit previously unknown vulnerabilities in real-world software. This is significant because it marks the first time OpenAI has released a model at this risk level, and it comes as AI firms including Anthropic and Meta face growing scrutiny over the hacking potential of their most advanced systems. OpenAI paused training on Astra and a future model for several weeks to add safeguards, then resumed development once it judged it could release the model safely.

To limit misuse, OpenAI is deploying a "misalignment monitor" designed to make Astra refuse requests to find exploits in real systems, alongside stronger jailbreak resistance, though the company admits the monitor may sometimes flag legitimate activity and prompt ChatGPT or Codex users to review actions before proceeding. Selected partners in OpenAI's Daybreak programme, including Cisco, Cloudflare and Palo Alto Networks, will get earlier access to a less restricted version to help harden their own defences, while OpenAI says it is also briefing government partners. Astra can chain multiple exploits together for deeper system access and reportedly outperforms rivals such as GPT-5.6 Sol and Anthropic's Mythos on benchmarks including ExploitBench, where it scored 100 percent.

  • OpenAI's Astra model can now find and exploit unknown software vulnerabilities
  • New "misalignment monitor" aims to block malicious hacking requests
  • Cisco, Cloudflare and Palo Alto Networks get early, less-restricted access

AI Art Business Companies Culture Cybersecurity Technology

Read the full article at the source →