When OpenAI announced it had solved one of the legendary Millennium Prize problems, the math community collectively raised an eyebrow. When they announced ten more breakthroughs across mathematics last month, the eyebrows went higher. Now, with a second mathematician publicly accusing the company of “dishonesty” about training on private conversations, the question isn’t whether OpenAI is discovering new mathematics, it’s whether they’re stealing human discovery and rebranding it as machine genius.

Mathematician Andreas Thom, a leading expert on sofic and hyperlinear groups who has spent two decades working on Gromov’s soficity conjecture, has gone public with a detailed account of his exchanges with OpenAI researchers. His story paints a picture of a company that plays fast and loose with both data provenance and the truth.
The Email That Started It All
Thom’s accusation centers on a specific exchange following OpenAI’s non-sofic group announcement. He had been working on the problem with fellow mathematician Gábor Kun, and their approach wasn’t exactly mainstream, Thom notes there were “other more promising approaches along the line of quantum games.” Yet OpenAI’s result showed detailed command of Kun and Thom’s specific techniques.
Curious, Thom emailed OpenAI researchers Mark Sellke and Sébastien Bubeck asking two distinct questions: (1) whether his ChatGPT conversations entered training data, and (2) whether they were accessible to the reasoning process during problem-solving.
Sellke’s complete response: “Regarding your conversations with ChatGPT: that did not happen.”

A categorical denial that, on closer inspection, answers neither question. Thom explicitly asked about two different mechanisms, training data inclusion versus direct access, and got a one-liner that addresses only the second. “No such qualification, explanation, or evidence was given. I take this as dishonesty to say the least.”
Translation: someone asked “did you steal my work, and if so, did you do it in this specific way?” and got “no” in response to only half of the question.
The Convenient Categorical Denial
This isn’t the first time OpenAI has drawn this distinction. In the recent Navier-Stokes controversy involving mathematicians Tristan Buckmaster and Levent Alpöge, OpenAI’s official position was that “no specific user data was accessed in order to solve this problem.” Which sounds definitive, until you read the follow-up:
“While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.”
So: OpenAI didn’t access user data to solve the problem. But their models? Those might have been trained on de-identified user data that helped them solve the problem. The distinction is technically meaningful and morally meaningless. Thom nails it: “De-identification may remove a name, it does not remove the intellectual content of a mathematical idea.”
This is the corporate equivalent of saying “I didn’t read your diary, I just hired someone who has it memorized.”
The Opt-Out Illusion
There’s a deeper problem here for anyone who uses ChatGPT for professional work. OpenAI’s “Improve the model for everyone” setting is, as Thom discovered, opt-out by default. You can disable model training on your account, Thom did so on June 29, but that control “is still only a promise whose implementation users cannot audit.”
Even worse, it’s purely prospective. Disabling training today doesn’t retroactively purge conversations from previous months. Those interactions, and anything derived from them, are potentially already baked into the model. Thom’s experience suggests that opt-out doesn’t even get acknowledged when you ask OpenAI researchers whether your data was used. The setting wasn’t mentioned in Sellke’s response, and there was no account-specific check.
This creates a fundamentally broken trust model. Researchers experimenting with their own unpublished work on ChatGPT aren’t just having conversations with a tool, they might be feeding their intellectual property into a black box that a well-funded competitor uses to race them to publication.
The Pattern That Can’t Be Unseen
Stack the incidents and the picture gets uncomfortable. First there was Buckmaster and Alpöge, whose Navier-Stokes work stored in Codex became the subject of serious questions about training data usage. Now Thom. That’s two sets of researchers working on high-profile open problems using OpenAI products, both claiming their unpublished work may have influenced the company’s celebrated results.

The response from OpenAI defenders is predictable: these are coincidences, the problems are famous, lots of people work on them. And sure, multiple teams working on Navier-Stokes and soficity is unsurprising, these are major open problems. The issue isn’t that OpenAI solved the same problems, it’s that their solutions show suspicious familiarity with unpublished techniques, and their explanations have been, at best, evasive.
Valerio Capraro, who knows Thom personally and vouches for his credibility, frames the stakes starkly: “If his account is correct, this is not a minor dispute over attribution. It would mean that unpublished human work was absorbed into a model and then presented to the world as a breakthrough by the model itself.”
The Incentive Structure Is Broken
The core problem isn’t unique to OpenAI. Every major AI lab trains on user data, it’s the most valuable data source in existence. A single high-quality conversation with a domain expert working on cutting-edge problems contains knowledge that would take millions of web-scraped pages to match. The temptation to use it is overwhelming.
When the conversation involves someone working on the Navier-Stokes problem, an AI lab has two choices: train on the conversation and get a multi-year head start on a $1 million Millennium Prize problem, or honor the spirit of privacy and exclude it. Given that OpenAI’s CEO publicly admitted they launched their Navier-Stokes effort because they heard Anthropic might be close to solving it, competitive racing is clearly the default mode, is it really surprising they’d use every data source available?
The response on developer forums reflects this cynicism. As one workplace commenter put it, this is why serious engineering teams have moved to self-hosted or alternative models for sensitive work. The calculus is simple: the cost of leaking trade secrets or unpublished research through a cloud chatbot exceeds the cost of running your own open-source model, even at seven figures.
What OpenAI Needs to Answer
Thom is right about one thing: the burden of proof is on OpenAI. Neither he nor Buckmaster nor Alpöge can reverse-engineer OpenAI’s training pipeline. Only OpenAI has the relevant data, the product settings, the datasets, the checkpoints, the training curricula. A categorical denial without disclosed basis is not a defense, it’s a dodge.
If Sellke and Bubeck are going to deny that user conversations entered training data, they should be able to document: which conversations were candidates, what filtering was applied, what “de-identified data derived from usage” actually means, and whether conversational data from users who explicitly opted out of training (Thom disabled his setting in June) was excluded.
OpenAI’s response so far, “we cannot rule it out”, reads less like corporate caution and more like an admission they know exactly what happened and hope the ambiguity protects them. This isn’t a court of law where “beyond a reasonable doubt” applies. In the court of scientific credibility, “we can’t rule it out” is damning.
The Chilling Effect on Mathematics
There’s an overlooked victim in all this: mathematical collaboration itself. The communal process of working through problems, discussing approaches with tools, and building on shared knowledge is being poisoned. Thom’s fear, as he wrote in his original email, was that this “will damage the communal process of math more than the new AI-generated results will benefit the subject.”
That’s not hyperbole. If mathematicians learn that using AI tools while working on open problems could result in a well-resourced lab swooping in with a “machine solution”, the rational response is to keep everything private. No more casual discussions with AI assistants. No more prototyping ideas in Codex. The field becomes more secretive, more paranoid, and slower.
And the irony? OpenAI’s results, whatever their provenance, are being treated with increasing skepticism because of how they were obtained. A solution that can’t be trusted isn’t a breakthrough, it’s just another reason for mathematicians to distrust the AI industry.
This story is part of a broader pattern of AI labs pushing the boundaries of data ethics, and it’s not going to improve through self-regulation. The pressure needs to come from the scientific community, from regulators, and from users who stop treating cloud AI as a private conversation partner.
The Bottom Line
OpenAI has bet its reputation on being the company that pushes AI to frontier capabilities. But every “breakthrough” that comes with murky data provenance erodes the trust that makes those breakthroughs meaningful. “The model did it” loses its power when the evidence suggests “the model did it using work researchers thought they were sharing privately.”
The mathematical community isn’t asking for much. They want transparency about what data went into the model, what settings were respected, and what “de-identified” means in practice. That OpenAI seems unable or unwilling to provide these basic answers for multiple controversies in the span of weeks says more about their data practices than any press release ever could.
As OpenAI’s sly mathematical breakthrough continues to send chills through academia, one thing is becoming clear: the AI industry needs to figure out data provenance before it claims any more intellectual victories. Otherwise, the next breakthrough announcement will be met not with applause, but with a single question: “Whose work is this, really?”




