Researchers used Claude to hack OpenAI employees’ ChatGPT accounts
Security researchers used Anthropic’s Claude models to exploit vulnerabilities in OpenAI’s community forum infrastructure, taking over several OpenAI employees’ ChatGPT and Codex accounts. They demonstrated the potential impact by instructing an employee’s Codex account to open a pull request in OpenAI’s internal repository, without accessing or altering internal code.
The exploit began with a heap buffer overflow in the libheif image-processing library, reached through HEIF uploads handled by Discourse and ImageMagick. Claude helped develop the exploit, which achieved remote code execution after a newer model was released; OpenAI fixed the issue within about 14 hours, paid a $6,500 bounty, and Discourse added image-processing sandboxing.
- Researchers compromised OpenAI accounts through an image-processing vulnerability.
- Claude helped develop the successful exploit.
- OpenAI fixed the flaw and paid a $6,500 bounty.