OpenAI's recent declaration about successfully addressing one of the celebrated Millennium Prize problems should ideally mark a significant achievement in the world of mathematics. While the outcome undeniably showcases the rapid advancements in AI's capabilities within mathematics, the details surrounding OpenAI's pursuit of the challenge have sparked controversy. Unprompted by its competitors, OpenAI seemingly mobilized its extensive resources in a last-minute drive to solve the problem first after hearing that others were making headway. This has led to serious accusations related to academic ethics, raising concerns that such behavior could stifle innovation within the field.
Mathematicians like Abhishek Saha from Queen Mary University of London have criticized OpenAI's approach, asserting that it strays from the conventions typically upheld by mathematicians.
In a blog post released on Tuesday, OpenAI revealed that it managed to derive a solution to the complex Navier-Stokes problem—a longstanding query concerning fluid dynamics—using one of its unreleased AI models, which achieved this in just 88 hours. This problem has long baffled mathematicians for nearly 90 years, and due to a $1 million reward associated with its resolution, it has attracted significant attention. OpenAI claims it achieved this by deploying a multitude of around 10,000 AI agents dedicated to the problem, designating the result a noteworthy milestone.
However, the timing of OpenAI's announcement raised questions. Just a day prior, NYU mathematics professor Tristan Buckmaster published research on a related topic alongside Levent Alpöge, a researcher from OpenAI’s rival, Anthropic. After learning that OpenAI had caught wind of their advancements, Buckmaster attempted to clarify when OpenAI had started addressing the problem and what training data its model utilized. This conversation reportedly devolved, with an OpenAI researcher suggesting that publicizing their findings could jeopardize Buckmaster’s career, infamously stating, "If you don’t want me to be nice, then I don’t have to be nice." OpenAI further urged him to publish his research while crediting OpenAI's internal model, excluding Alpöge as a co-author.
Buckmaster expressed concerns that OpenAI may have accessed his sessions on Codex, the platform he used to work on the problem. However, OpenAI has consistently denied any acquisition of specific user data, asserting that its solutions were developed independently and prior to the public release of Buckmaster's work. Yet, the company did not entirely dismiss the possibility that anonymized data from their products might have contributed to their model's development, but emphasized that the two approaches to proofs are distinct.
The exact nature of events remains unclear, as the timelines are convoluted and the overlap in research makes tracing the origins of AI-generated ideas challenging. It is predictable that various entities would vie to resolve one of math's most esteemed challenges, especially given the sizable reward attached.
Nevertheless, elements of OpenAI's narrative appear puzzling. In its own description, the endeavor was a rushed and costly effort, reportedly consuming millions of dollars. Yet OpenAI has stated it does not seek the prize, which is still pending from the Clay Mathematics Institute. They claim their motivation lies solely in illustrating the significant progress of their AI models and did not seem to have extensively explored Navier-Stokes prior to September.
OpenAI’s rationale for its expedited efforts seems to boil down to the question: "Why not?" According to OpenAI researcher Sébastien Bubeck, the company became aware of rumors about competing researchers tackling Millennium Prize problems via Twitter and decided to explore this path, believing their model was sufficiently robust.
Buckmaster's allegations regarding the misuse of his data remain mostly unaddressed by OpenAI aside from refuting the specific claims. Bubeck has disputed some of Buckmaster's account, claiming he never suggested he be removed as a co-author.
Even beyond the more serious accusations, OpenAI’s actions reflect a broader discontent among mathematicians. The rush to claim a breakthrough contradicts the collaborative spirit typically seen in the mathematical community. Saha notes that the nature of advanced research requires profound specialization that normally prevents unjustified last-minute interventions, while Matthew Ballard from the University of South Carolina emphasizes the importance of a trusting environment in mathematics, where researchers traditionally exchange incomplete ideas without the fear of competition.
Concerned about how this environment might shift, Carnegie Mellon professor Jeremy Avigad voiced unease that AI systems potentially infringing on researchers' intellectual property could make mathematicians more reserved in sharing ideas. The sight of a vast number of AI agents converging on an unresolved challenge may foster a culture of caution, which would alter the collaborative nature of research.
OpenAI's ambiguity regarding whether its models may have been influenced by existing mathematical work heightens anxiety in the mathematical community. Brendan Hassett from Brown University points out that, given the established history of AI companies utilizing material without permission, skepticism among researchers is expected. Companies should be held accountable to ensure the protection of researchers' contributions, he asserts.
Looking ahead, it remains uncertain how the landscape will transform. Yang-Hui He, a fellow at the London Institute for Mathematical Sciences, shares concerns about the growing secrecy surrounding corporate involvement in mathematics, reminiscent of past eras dependent on wealthy benefactors.
For many researchers, the future may not undergo radical changes, as there exists a finite number of high-profile problems organizations like OpenAI would realistically pursue. Saha suggests that AI labs may not choose to allocate vast resources to most ongoing mathematical inquiries, as the payoff in exposure might not be sufficient.
Some speculate that the race to tackle this problem may have been influenced by the potential for significant media attention, as suggested by Oxford’s Andras Juhasz, indicating this might have been a strategic PR win for OpenAI. However, the sustainability of this approach could be in question, especially given that the ability of AI models to address complex questions means that the scale of competition could escalate. As Juhasz points out, human mathematicians also face the risk of being ousted by one another; however, the scale at which OpenAI executed this could be uniquely daunting.
While OpenAI has showcased how its capabilities can directly compete in advanced mathematics, the reaction from the professional community suggests that it may have alienated the very segment it aimed to engage.




