← Back to the feed

OpenAI’s hundreds of mathematical results may take years to assess

The Verge ·

OpenAI has released nearly 400 AI-generated mathematical results across more than 700 manuscripts, spanning fields from number theory to mathematical physics. Researchers described the scale as overwhelming and said understanding the work, let alone assessing its impact on mathematics, could take years.

The results are at varying stages of verification: OpenAI said 300 of 719 headline results, about 42 per cent, had been formalised in Lean, a proof-checking system. Mathematicians cautioned that formalisation still needs scrutiny and that some proofs may not clearly match the claims in their papers; others worry the volume could include errors and that further releases may arrive before the current work is assessed.

  • OpenAI released nearly 400 AI-generated mathematical results.
  • Only about 42 per cent of headline results were formalised in Lean.
  • Researchers say assessing the collection could take years.

New here? Start with this

OpenAI, an artificial intelligence company, has released nearly 400 mathematical results generated by its AI systems. These results span more than 700 academic papers across various mathematical fields, including number theory and mathematical physics. This approach differs from how mathematics has traditionally developed, where individual mathematicians or small teams typically work through problems over extended periods.

For the mathematical community, verifying whether these results are correct is important. About 42 per cent have been formally checked using a specialised computer system called Lean that verifies mathematical proofs. However, some mathematicians question whether this verification is sufficient; some proofs may not clearly support the claims in the papers, and errors could be present.

The large volume of results creates a significant challenge for assessment. Experts say that understanding and evaluating these findings could take years, and some worry that additional releases may come before the current batch has been thoroughly examined.

Both sides, in good faith

The strongest fair case each way — we don't pick a winner.

The case for

Those supporting the release argue that the sheer volume of potential mathematical insights is inherently valuable, particularly given that 42 per cent have been formally verified in Lean, and that transparency serves the mathematical community well. They contend that distributed peer review across researchers accelerates discovery in ways that withholding unverified work cannot, and that the alternative—keeping such results private until perfect verification—harms a field that could benefit from computational exploration and open scrutiny.

The case against

Those concerned about the approach contend that releasing hundreds of unverified results, with 58 per cent lacking formal proof-checking, risks flooding mathematics with potentially erroneous claims that consume researchers' finite time and attention. They argue that the stated overwhelm and volume indicate the community cannot reasonably assess the work, which contradicts scientific norms requiring verification before broad dissemination, and that successive releases before current work is evaluated prioritise demonstrating AI capability over genuine contribution to mathematical knowledge.

AI Technology

Read the full article at the source →

Originally published by The Verge as “‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest drop”.