OpenAI’s Navier–Stokes Claim Rekindles Debate Over AI and Mathematical Insight
Summary
Michael Li traces a year of AI-assisted mathematical results leading to OpenAI’s September 8 announcement that an internal model had produced a proof of finite-time blowup for the forced three-dimensional Navier–Stokes equations, a problem associated with a Millennium Prize. OpenAI said the effort used 10,000 agents for 88 hours and produced both a proof and a Lean formalization. The article presents this as a sharp departure from the May result on the Erdős unit distance conjecture, which was accompanied by a paper from nine mathematicians who explained and simplified the argument. During the summer, other results attributed to AI or AI-assisted work included a counterexample to the Jacobian conjecture, a resolution of the Hadamard matrix conjecture, a complex structure on the six-sphere, a non-sofic group construction, and an increase in the reported lower bound for zeros of the zeta function on the critical line from 41.6% to 67.2%. Several of these announcements lacked companion papers that would make the underlying ideas accessible to mathematicians, even when Lean certificates or independently checkable examples supported correctness. Terence Tao and Hugo Duminil-Copin criticized the use of mathematics as a source of AI hype and warned that automated systems could produce answers faster than fields can develop understanding. Tao specifically argued that undisclosed automated solutions could contaminate problems as benchmarks and discourage researchers from sharing promising directions. The Navier–Stokes episode intensified those concerns: mathematicians Tristan Buckmaster and Levent Alpöge had posted related results shortly before OpenAI’s announcement, and Buckmaster said OpenAI did not directly answer questions about when its model was prompted or whether data from their Codex work could have influenced training. OpenAI said its researchers had not seen their work before public release but could not rule out de-identified product data having improved its models. OpenAI also acknowledged that its team lacked research-level expertise in fluid dynamics. The article concludes that the rush to verify and announce a result may have produced a technical answer without the broader mathematical account needed to turn it into shared insight, while the proof’s ultimate status remains tied to expert scrutiny.