OpenAI Says 10,000 AI Agents Cracked a $1M Math Problem. Then a Fight Broke Out.

OpenAI says one of its unreleased models just solved a problem that has beaten the finest human mathematicians for over a century, and then, almost on cue, a very human fight broke out about who really did it. On 8 September the company announced that a next-generation system, described as significantly more capable than GPT-6 Astra, had produced a proof for the Navier-Stokes existence and smoothness problem, one of the seven Millennium Prize Problems, each worth a cool $1 million. It reportedly took roughly 10,000 AI agents working in parallel for 88 hours.

If it holds up, it is a genuinely historic moment: the first time a machine has cracked a problem at the absolute frontier of mathematics rather than merely helping a human do it. That “if” is doing an enormous amount of work, so let us take the claim seriously and sceptically at the same time, which is the only sensible way to treat anything a frontier lab announces about its own unreleased model.

What the proof actually claims

The Navier-Stokes equations describe how fluids move: water, air, weather, blood. The Millennium question asks, roughly, whether the equations always behave nicely in three dimensions, or whether they can “blow up”, producing an impossible infinity, and thus failing to describe reality. OpenAI’s model reportedly landed on the dramatic answer: it described a configuration where smooth fluid motion breaks down, a vortex tightening and spinning ever faster in what mathematicians call finite-time blowup, while the fluid’s energy stays bounded throughout.

Crucially, the company did not just post an argument and ask for trust. It released a paper and a proof written in Lean, a formal theorem-proving language that lets a computer verify each logical step mechanically. That matters. A Lean-checked proof is far harder to wave away than a wall of natural-language mathematics, and it is the single strongest point in OpenAI’s favour here. Verification by the wider mathematical community will still take time, and until that is done the honest label is “claimed”, not “confirmed”.

Then the row started

Here is where it gets human. According to reporting picked up by Quanta and others, OpenAI says its effort began on 1 September after it heard a rumour connected to two researchers: Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a maths professor at NYU. OpenAI finished its own proof and Lean verification by 6 September, then contacted the pair to offer a joint announcement, at which point, it says, it discovered their work actually addressed a related but distinct problem, the forced Euler equations rather than Navier-Stokes.

Buckmaster tells it differently. He alleges OpenAI pressured him to exclude Alpöge from authorship after learning of their progress. OpenAI’s researchers deny it flatly, saying they never accessed the pair’s private work or their Codex logs. We are in he-said territory and this piece is not going to pretend to know who is right. But the dispute is revealing regardless of who has the better case, because it is a preview of a question the field has not begun to answer: when a machine and some humans converge on the same mathematical frontier at the same moment, who gets the credit, and who decides?

Why this connects to everything else this week

Step back and the timing is almost too neat. In the same stretch of days that an Anthropic researcher quit warning that self-improving AI is racing out of control, OpenAI announced that a swarm of 10,000 coordinating agents, built on a model it has not released, autonomously solved a problem humans could not. Whatever you make of the extinction debate, this is a concrete data point about capability: not a chatbot writing your emails, but a system doing original research at the edge of human knowledge.

Keep both the wonder and the caution. The wonder is obvious and earned. The caution is that “our secret model did something amazing, trust us” is exactly the kind of claim that deserves independent scrutiny before the champagne, and the attribution mess is a reminder that the humans around these systems still have all the usual incentives to grab credit. If the proof survives peer review, it is a landmark. If it does not, it will be a very expensive press release. Either way, a machine is now a serious participant in the argument, and that alone is worth marking.

Related on Top Tool Stack: the Anthropic researcher who quit over self-improving AI and OpenAI’s GPT-6 Astra launch.

Get the free weekly stack: the AI tools and moves that matter, hype filtered out.Subscribe free →
Scroll to Top