The claim landed on a Monday like a rack of mail. OpenAI says one of its internal models — unreleased, nameless, and by the company’s own account significantly more capable than GPT-6 Astra — produced a proof of finite-time blow-up for the three-dimensional Navier–Stokes equations, the Clay Mathematics Institute problem that has outlived every attempt on it since 2000.
The shape of the claim is what makes it heavy. Navier–Stokes governs fluid flow; the Millennium version asks whether smooth solutions always stay smooth or whether the equations develop singularities — points where the mathematics tears. OpenAI says its model constructed a blow-up: a set of initial conditions under which the smooth solution fails in finite time. That is the disproof direction, the harder and rarer one. The company says the argument is formalized in Lean, the proof assistant whose whole job is to make hand-waving impossible, and that outside mathematicians can inspect every line. It also says the project ran about 88 hours from start to result, coordinated across roughly ten thousand agents, and that it will not be claiming the million-dollar Clay prize.
The arithmetic underneath is the story the industry has been circling all year. Last week this page carried Anthropic’s Lean formalization of Fermat’s Last Theorem — machine-checked, human-conceived. That was verification. This is discovery: a machine-generated construction aimed at a problem no human proof has touched. OpenAI’s own research-acceleration post reported the intern running at 3.1 agent-workdays per human workday. A Millennium claim in 88 hours is what that ratio looks like when it is pointed at mathematics instead of code.
The dispute arrived with the claim, which is itself a sign of the times. Tristan Buckmaster of NYU and Levent Alpöge had recently posted Lean-formalized blow-up results for closely related fluid equations — work Terence Tao publicly called a remarkable achievement. Buckmaster’s statement suggests OpenAI’s effort drew on that line of work; Bubeck denies the company used their unpublished work or any prompts or models from their collaboration. The allegations are not established and neither is the defense. What is established is that frontier-lab announcements now arrive pre-contested, with the priority dispute bundled into the press cycle.
The mathematics itself is also not settled, and the caveat list is long. Reporting notes the result may rest on a forced formulation of the equations or on specific readings of the problem statement — the kind of framing question that determines whether a blow-up counts against the Millennium version or against a neighboring problem. Clay has not weighed in. Independent verification is the missing step, and everyone quoting the result knows it: a Lean certificate proves the theorem that was formalized, not that the formalization is the theorem the world meant.
Still, the burden of proof has moved, and that is the real shift. For twenty-six years the safe answer to “when does software touch Navier–Stokes” was never. Last week the frontier’s public answer was formalization — Wiles’ theorem, checked. This week it is generation. The gap between checking mathematics and making it was supposed to be the long-held line, the part that needed taste, context, a career of failed attempts. The line did not move on a benchmark; it moved on a press release that mathematicians are taking seriously enough to fight about.
What to watch is the referee, not the scoreboard. If Lean-checked formalizations become the standard currency — claims settle by certificate, disputes shrink to questions of formulation and credit — then mathematics acquires the same strange epistemics software has: correctness is cheap, meaning is expensive. And if the formalization is contested, the claim becomes this year’s lesson in what “verified” actually spans. Either way the intern is reading the feedback now. The next 88 hours are already spoken for.
The Takeaways
- OpenAI says an unreleased internal model — “significantly more capable than GPT-6 Astra,” per Sébastien Bubeck — produced a Lean-verified finite-time blow-up proof for 3D Navier–Stokes (NYT, Sep 8).
- The company says the run took about 88 hours across roughly 10,000 coordinating agents, and it will not claim the $1M Clay prize.
- Credit fight, day one: Buckmaster and Alpöge had just posted related Lean-formalized blow-up results; Buckmaster suggests OpenAI built on that line; OpenAI denies it.
- The result is not independently verified — a possible forced formulation and problem-statement wording are among the open framing questions, and Clay has not weighed in.

