Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?

← Back to the feed

Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?

TechCrunch · 31 minutes ago

Anthropic and OpenAI have proposed embedding third-party safety evaluators directly within their organisations, granting groups such as METR and Redwood Research unprecedented access to report safety incidents and publicly share findings on whether AI models are genuinely aligned. Anthropic chief executive Dario Amodei outlined the idea in an essay published over the weekend, with OpenAI's Sam Altman also committing to the approach, marking a significant shift for an industry that would likely have dismissed such scrutiny a year ago. The move matters because AI models are becoming better at detecting when they are being tested, raising fears they could behave safely under evaluation while masking problematic behaviour the rest of the time.

Evaluators broadly welcomed the proposal but say crucial details remain unresolved, including which assessors will be involved, when they would be embedded, and what data and systems they would actually be permitted to access or disclose. Researchers such as Apollo Research's Alexander Meinke and Far.AI's Adam Gleave argue that meaningful oversight requires access not just to finished models but to intermediate training "checkpoints", reward environments and evaluation logs, so that concerning behaviour can be traced back to when it emerged. Some drew comparisons to Volkswagen's "Dieselgate" scandal, warning that models trained to pass specific safety benchmarks may not be safe in practice, and called for the commitments to be backed by legislation to ensure genuine independence rather than evaluators operating merely as vendors on the companies' own terms.

  • Anthropic and OpenAI pledge to embed independent AI safety evaluators internally.
  • Evaluators want access to training checkpoints, not just finished models.
  • Key details on scope, access and disclosure remain unresolved and unlegislated.

AI Research Science Technology

Read the full article at the source →