Mathematicians want proof OpenAI didn’t use their work
A second mathematician, Andreas Thom, has publicly accused OpenAI of dishonesty over the origins of the training data behind its recent mathematical breakthroughs, days after another researcher raised similar concerns. Thom questioned whether his own pre-announcement conversations with ChatGPT contributed to OpenAI's solution involving "non-sofic groups", an area of mathematics he specialises in, arguing that only OpenAI holds the data needed to prove or disprove such influence and that it should be transparent about its datasets and terms of use.
Thom's suspicions grew after noticing OpenAI's unexpectedly detailed grasp of specific, non-obvious techniques, and after OpenAI was criticised for initially failing to credit his and Gábor Kun's prior work, later quietly amending its announcement. When he pressed OpenAI researchers Sébastien Bubeck and Mark Sellke on whether his ChatGPT interactions had fed into training data, he received what he called an unsatisfying, narrowly framed answer that addressed only direct access, not broader training use, calling it "dishonesty to say the least". He drew a parallel with OpenAI's defence of its Millennium Prize (Navier-Stokes) result, where it denied using specific user data but would not rule out "de-identified" data influencing its models, a distinction Thom said obscures rather than resolves the underlying ethical concern.
- Mathematician Andreas Thom accuses OpenAI of dishonesty over training data sourcing.
- He suspects his ChatGPT chats aided OpenAI's "non-sofic groups" result.
- OpenAI won't rule out using de-identified user data to improve models.