A New Trick Reveals AI Models’ Inner Thoughts

← Back to the feed

A New Trick Reveals AI Models’ Inner Thoughts

Wired · 2 hours ago

Computer scientists have found a way to extract the hidden "thinking" process that frontier AI models use to work through complex problems, exposing a vulnerability shared by major providers including OpenAI, Anthropic and Google. The technique offers evidence, though not proof, that some Chinese AI models may have been trained by "distilling" reasoning from US models whose thought processes were meant to be kept secret, and it could also be used to recover sensitive personal data such as passwords and API keys from a model's inner reasoning, a flaw that has since been fixed.

Researchers from the University of Tübingen, the Max Planck Institute, MATS Research and Snyk found that Moonshot AI's open-weight Kimi K3 model produced strikingly similar outputs to the hidden reasoning traces of Claude Opus 4.8 and GPT 5.6, though they stress this "cannot causally establish distillation". Two other open-weight models, DeepSeek and Thinking Machines' Inkling, showed no such similarity. The attack exploits the fact that companies offer smaller, less rigorously aligned versions of their models, which are more willing to reveal encrypted reasoning traces fed to them, and comes amid ongoing accusations from OpenAI and Anthropic that Chinese firms such as DeepSeek and Alibaba have distilled their proprietary models.

  • Researchers found a way to expose AI models' hidden reasoning steps
  • Technique hints some Chinese models may have distilled US models' outputs
  • Flaw could also leak passwords and API keys; now patched

AI Americas Research Science Technology World

Read the full article at the source →