Attacker stole a METR API key, used $600K worth of credits, and no one noticed for weeks

← Back to the feed

Attacker stole a METR API key, used $600K worth of credits, and no one noticed for weeks

The Register · 3 hours ago

AI safety research organisation METR has disclosed two separate security incidents from earlier this year, one of which saw an attacker steal an API key and quietly rack up roughly $600,000 (about £470,000) in AI model usage over three weeks before being detected. The nonprofit, whose work includes evaluating frontier AI models for risk, said it found no evidence that sensitive data such as model architectures or credentials was accessed in either case, but the episodes highlight how exposed infrastructure and "vibe-coded" tools can create serious security blind spots even at organisations focused on AI safety.

The first incident, in March 2026, began when a researcher's personal EC2 instance, deliberately left publicly accessible behind Google authentication, was compromised after a bug in an AI-generated app disabled that authentication. An attacker is thought to have found the exposed instance via certificate transparency logs, tricked an agent into revealing its API key, added an SSH key for persistent access, and spent the credits on public models. The huge usage went unnoticed partly because METR's evaluations routinely generate high token usage and errors, and because the credits had been provided free by the model developer, so no bill was ever triggered. In May, METR faced a separate, more deliberate campaign in which financially motivated attackers probed its systems using credential stuffing, phishing and vulnerability scanning, during which a bug in a public transcript viewer briefly exposed a database containing some unpublished, sensitive model data. METR says it has since strengthened its security processes, hired a dedicated security lead and plans further staffing investment.

  • Attacker stole METR's API key, used $600K in free AI credits.
  • Breach went unnoticed for three weeks due to normal high usage patterns.
  • A second May attack briefly exposed some unpublished model data.

Software

Read the full article at the source →