OpenAI pledges to add Astra security as Anthropic loosens Fable’s leash

← Back to the feed

OpenAI pledges to add Astra security as Anthropic loosens Fable’s leash

The Register · 7 hours ago

OpenAI says its forthcoming Astra model may have cyber capabilities severe enough to create new forms of harm, despite not being involved in the recent Hugging Face incident. The company has pledged tighter safeguards during development and testing, highlighting wider concerns over whether frontier AI systems are being adequately secured before release.

The proposed measures include isolated test environments, restricted network and tool access, stronger protection of model weights, monitoring and sandboxed execution; OpenAI says it will pause internal testing where these controls are missing. It is also monitoring Astra’s chain-of-thought during pre-release work for risky actions, while Anthropic has said it will reduce Fable’s refusals for some biology-related prompts after criticism that its earlier restrictions hindered legitimate researchers.

  • OpenAI plans stricter safeguards for its Astra model.
  • Astra may possess potentially critical cyber capabilities.
  • Anthropic is easing some biology-related Fable restrictions.

AI Americas Technology World

Read the full article at the source →