Another researcher is challenging OpenAI about the data that might be fueling its impressive mathematical breakthroughs. Just days after a heated debate over whether the company’s models benefited from unpublished work, a second mathematician has accused OpenAI of unethical and "dishonest" behavior, citing a lack of transparency about the origins of its training data. Mathematicians are demanding proof that OpenAI did not use their research without permission. Another mathematician accused the company of "dishonesty" after a series of major AI-driven discoveries. Mathematician Andreas Thom raised concerns in a series of posts on Mastodon, suggesting that his and his colleagues' interactions with the ChatGPT chatbot before OpenAI’s recent announcement may have contributed to the AI’s success in mathematics. One of the 10 results OpenAI celebrated last month involved non-sofic groups, an area of mathematics where Thom is an expert. OpenAI acknowledged that their result built on previous work by Thom and Gábor Kun, but the company faced criticism for not fully recognizing their contributions. Afterward, OpenAI quietly revised its announcement to include more credit. Non-sofic groups are, roughly speaking, infinite mathematical structures that cannot be approximated by finite ones. Thom was particularly struck by "OpenAI’s detailed command of our techniques," which he said were neither the most obvious nor the most promising routes to a solution. He wrote to OpenAI researchers Sébastien Bubeck and Mark Sellke, asking whether his interactions with ChatGPT were part of the training data or accessible to the reasoning process. However, the response he received did not satisfy him, as it only addressed whether his conversations could be accessed directly, not whether they had entered OpenAI’s vast training data. "No such qualification, explanation, or evidence was given," he wrote. "I take this as dishonesty to say the least." Thom emphasized that researchers are not equipped to reverse-engineer OpenAI’s training pipeline to determine if their work has been used. "Only OpenAI has the relevant data for that," he said. If the company is to deny using user data, he argued, it must prove it by disclosing all necessary datasets and clarifying how it uses data. OpenAI’s reluctance to rule out the use of user data echoes how it defended its recent Millennium Prize breakthrough, both in public messaging and in communications with Tristan Buckmaster, a mathematician who questioned whether OpenAI’s models used his work. In its blog post announcing the Navier-Stokes solution, which concerns the movement of fluids, OpenAI denied using any specific user data. However, it did not conclusively rule out an indirect influence, stating that "de-identified data derived from their usage of our products helped improve our models" could not be ruled out. Thom said this is the same distinction OpenAI used in its communications with him. "De-identification may remove a name; it does not remove the intellectual content of a mathematical idea," he said. Thom criticized Sellke’s response as "unjustifiably broad and materially misleading," adding that it was "plainly dishonest" to use nonpublic research without consent or proper credit. He argued that it would be ethically indefensible if nonpublic research helped improve models that the company then used to race users to publication. OpenAI has not yet responded to The Verge’s request for comment. His comments reflect growing unease in the mathematical community, as OpenAI’s recent achievement in solving a legendary Millennium Prize problem is overshadowed by concerns about data use and ethical practices. The incident has left a sour taste among researchers, who worry that such behavior could lead to a more secretive academic environment if mathematicians fear their work could be exploited by powerful tech companies.