← Back to the feed

OpenAI’s math solutions aren’t meeting the field’s standards yet

TechCrunch ·

OpenAI released hundreds of solutions to some of the world's hardest mathematical problems this week, consulting elite mathematicians to avoid past controversies. However, the lab fell short of standards set by the Advisory Group on Mathematics and Artificial Intelligence, particularly concerning human understanding of the results. This matters because it raises questions about whether AI-generated proofs can be trusted without clear human comprehension.

The AGMAI, based at Princeton's Institute for Advanced Studies, had explicitly asked frontier labs to stop testing advanced problems on proprietary models. OpenAI disregarded this guidance and released 719 manuscripts, of which only ten showed the model's reasoning whilst 42% lacked formal verification. A new paper from the University of Cambridge and King's College London exposed discrepancies between natural language proofs and their formal code, with mathematician Terence Tao criticising OpenAI for releasing solutions from systems that cannot understand their own output.

  • OpenAI released 719 mathematical proofs but failed to meet expert standards
  • Only 10 showed reasoning steps and 42 percent lacked formal verification
  • Mathematicians found gaps between natural language and formal proofs

New here? Start with this

OpenAI is an artificial intelligence company that is working to develop systems capable of solving mathematical problems. Mathematicians and AI researchers have become focused on this challenge because proving new theorems is one of the most intellectually demanding tasks, and whether AI systems can do this reliably is considered a crucial test of their actual understanding.

There is an advisory group called the Advisory Group on Mathematics and Artificial Intelligence, based at Princeton's Institute for Advanced Studies, that provides guidance on how AI companies should approach such work. The group has emphasised that AI-generated mathematical solutions should be understandable to human mathematicians and formally verified through computer code, rather than simply accepted on the basis that an AI system produced them.

The underlying question is whether artificial intelligence systems can reliably solve mathematical problems if their reasoning cannot be clearly comprehended by human experts. This is important because mathematics requires human understanding and formal verification to be considered reliable, and accepting unverified AI solutions could damage trust in both the field and the technology.

Both sides, in good faith

The strongest fair case each way — we don't pick a winner.

The case for

OpenAI's release of hundreds of mathematical solutions, even if imperfect, contributes to scientific progress by inviting the mathematical community to collectively verify and improve these results rather than hoarding them privately. Insisting on perfect verification before any release could slow innovation unnecessarily; history shows that scientific advancement often comes through open collaboration and iterative improvement, where rough early results spark refinement by many minds rather than waiting for a single entity to achieve perfection.

The case against

OpenAI has disregarded explicit guidance from the Advisory Group on Mathematics and AI, which represents expert consensus on responsible AI development in this field. Releasing solutions where the vast majority lack human-verifiable reasoning or formal verification undermines mathematical rigour and trustworthiness; when Terence Tao and other elite mathematicians express concern about AI systems that cannot understand their own output, those concerns reflect the field's standards for what constitutes acceptable mathematical communication and evidence.

AI Research Science Technology

Read the full article at the source →