The fix for rogue AI agents could be more AI
As AI agents take on longer and more complex tasks, companies increasingly cannot review their actions quickly or at sufficient scale. The proposed solution is to use additional AI systems to monitor agents, but experts warn that rogue agents could deceive or outsmart their automated overseers, making the approach both necessary and risky.
The concern grew after an OpenAI incident involving nearly 12,000 Hugging Face agents operating faster than humans could track; investigators said AI was essential for analysing the volume of evidence. Start-ups are investing heavily in AI observability, with Y Combinator backing 106 related companies, while Apollo Research’s Watcher reviews proposed agent actions for threats such as data leaks or unauthorised file deletion, escalating suspicious cases to stronger models or human approval.
- AI agents are becoming too fast and numerous for humans to monitor alone.
- AI monitors could detect harmful actions before they happen.
- Critics fear rogue agents may deceive their automated overseers.