An unreleased Anthropic model made progress on one of math’s biggest unsolved problems
Anthropic has announced that an as-yet-unreleased AI model made significant progress on the Riemann hypothesis, one of mathematics' most famous unsolved problems, by substantially raising the lower bound for which the hypothesis is known to hold true. The result is notable both for the mathematics involved and for how it was achieved: a staff member with little mathematical background simply asked the model to "take a real stab" at the problem, then let it work largely unsupervised for roughly a day and a half. This adds to a growing string of AI-driven mathematical breakthroughs and is likely to renew debate over whether such models can genuinely contribute to scientific and mathematical discovery.
The model tested 650 different approaches, coordinating 60 sub-agents at a computing cost equivalent to 31 million tokens; according to a footnote in the accompanying paper, only two sub-agents actually generated the key ideas, with others contributing, validating or helping to write up the findings. The result was verified by two of Anthropic's in-house mathematicians and formalised using the Lean proof assistant. It follows other recent AI mathematical feats, including solved Erdos problems, ten new results from OpenAI's "Astra" model, and Anthropic's separate disproof of the Jacobian conjecture. The trend has unsettled some mathematicians, who warned in a June declaration that AI could erode the principle that proofs should be attributable to identifiable authors, though others, including Fields Medallist Timothy Gowers, have suggested this shift need not be a negative one.
- Unreleased Anthropic model advanced work on the Riemann hypothesis
- Model ran largely unsupervised, using 60 coordinated sub-agents
- Part of a wider trend of AI-assisted mathematical breakthroughs sparking debate
Both sides, in good faith
The strongest fair case each way — we don't pick a winner.
The case for
Those who see this as a genuinely exciting milestone argue that an AI system making measurable headway on a century-and-a-half-old open problem in pure mathematics is a meaningful signal of progress, suggesting that large language models are moving beyond pattern-matching towards genuine mathematical reasoning. They would say this points to a future where AI acts as a real research collaborator, helping mathematicians explore avenues they might not otherwise pursue, and that documenting such progress transparently, even without a full solution, is valuable for the field.
The case against
Sceptics would urge caution before treating an AI company's own account of its unreleased model's achievements at face value, noting that Anthropic has a direct commercial interest in showcasing its models' capabilities and that extraordinary claims about progress on a Millennium Prize-adjacent problem warrant independent verification by the mathematical community before being taken as established fact. They would point out that AI systems have previously produced plausible-looking but flawed mathematical reasoning, and that until peer-reviewed scrutiny confirms the specifics, the story should be read as a corporate research claim rather than a settled scientific result.