Mathematicians want proof OpenAI didn’t use their work 

← Back to the feed

Mathematicians want proof OpenAI didn’t use their work 

The Verge · 2 hours ago

A second mathematician, Andreas Thom, has publicly accused OpenAI of dishonesty over the origins of the training data behind its recent mathematical breakthroughs, days after another researcher raised similar concerns. Thom questioned whether his own pre-announcement conversations with ChatGPT contributed to OpenAI's solution involving "non-sofic groups", an area of mathematics he specialises in, arguing that only OpenAI holds the data needed to prove or disprove such influence and that it should be transparent about its datasets and terms of use.

Thom's suspicions grew after noticing OpenAI's unexpectedly detailed grasp of specific, non-obvious techniques, and after OpenAI was criticised for initially failing to credit his and Gábor Kun's prior work, later quietly amending its announcement. When he pressed OpenAI researchers Sébastien Bubeck and Mark Sellke on whether his ChatGPT interactions had fed into training data, he received what he called an unsatisfying, narrowly framed answer that addressed only direct access, not broader training use, calling it "dishonesty to say the least". He drew a parallel with OpenAI's defence of its Millennium Prize (Navier-Stokes) result, where it denied using specific user data but would not rule out "de-identified" data influencing its models, a distinction Thom said obscures rather than resolves the underlying ethical concern.

  • Mathematician Andreas Thom accuses OpenAI of dishonesty over training data sourcing.
  • He suspects his ChatGPT chats aided OpenAI's "non-sofic groups" result.
  • OpenAI won't rule out using de-identified user data to improve models.

New here? Start with this

Suspicion has grown that leading AI companies, including OpenAI, may have trained their systems on private or semi-private exchanges with researchers, then used the results to claim credit for solving hard problems. The row centres on mathematics: OpenAI said its AI had made breakthroughs on tricky, unsolved-style problems, but some mathematicians whose earlier work touches on the same territory are now asking how the AI got so good, so fast, and whether their own ideas ended up feeding it without proper acknowledgement or consent.

The key figures are OpenAI, the company behind ChatGPT, and independent mathematicians, including Andreas Thom, who study niche areas of the field. Thom had exchanged messages with ChatGPT before OpenAI's announcement, and later noticed the tool's answers reflected an unusually precise understanding of his own specialist techniques. OpenAI staff have responded to questions from mathematicians about this, but not in a way that has settled the dispute.

At the heart of the matter is a broader question about how AI companies gather the data used to train their models, and how transparent they are about it. Researchers argue that only the companies themselves can say for certain whether private conversations, papers or unpublished work end up shaping an AI's output, which makes it hard for outsiders to verify claims of originality. This matters because it touches on trust in AI-generated research, proper credit for human scientists, and how openly tech firms disclose what their systems have actually learned from.

AI Business Companies Technology

Read the full article at the source →