A scientist questions the transparency of OpenAI's data
Andreas Thom has expressed doubts about the transparency of OpenAI's use of user data following the similarity of methods used to solve a complex problem. He is urging the company to disclose more details, but OpenAI has not made any public comments on the situation.
Crius
Andreas Thom, a group theory specialist from the Technical University of Dresden, has expressed doubts about whether his personal ChatGPT sessions were used to achieve OpenAI’s result related to the construction of the first known non-sofic group—one of ten problems that, according to OpenAI, were solved by Astra on August 1.
In a post on Mathstodon, Thom shared that shortly after the announcement, he sent an email to OpenAI researchers Mark Selke and Sébastien Bubeck. Thom and his colleague from Dresden had spent several months working on the problem of comparing expanders and extensions, building on their previous work with Gabor Kun, using ChatGPT. In his email, he asked two questions: whether these exchanges had been included in the training data, and whether the system could have accessed them while working on the proof. In response, Selke stated that this had not happened. Thom notes that this answer addressed only the issue of “direct access” and considers it too broad and, in retrospect, inaccurate.
This exchange has resurfaced for discussion this week following OpenAI’s announcement about Navier-Stokes. In OpenAI’s statement, it is noted that neither the researchers nor the company’s agents had seen the work of this pair before its publication, and no specific user data was used. It was also added that the possibility of using anonymized data obtained through the use of the company’s products to improve models cannot be ruled out.
Thom offers two more arguments. He disabled model training on June 29, but users cannot verify this setting, and it does not affect earlier chats. Moreover, most researchers preferred quantum game methods rather than the approach used by Thom and Kun, so he was surprised that Astra chose the same path.
At present, Thom refrains from making claims about the existence of evidence and urges OpenAI to disclose the basis for its denial. OpenAI has not publicly commented on his posts.
