10,000 agents, 130 billion tokens — and less than 10 percent of the credit for the swarm
When OpenAI announced last week that a system of roughly 10,000 AI agents had used 130 billion tokens over 88 hours on a Millennium Prize problem, it was read as a breakthrough for multi-agent AI. But when researcher Noam Brown appeared on the Dwarkesh podcast on 17 September 2026, he lowered expectations himself: according to a summary of the episode, he attributes less than 10 percent of the credit to multi-agent coordination. The rest is due to a generally capable base model that manages to work over long time horizons. It is a rare honest look at what one of the year's most talked-about AI results actually demonstrates — and what it does not.
The news: the swarm on the Millennium Prize problem
Host Dwarkesh Patel opens the episode by saying that OpenAI last week announced it had solved one of the Millennium Prize problems with a system of 10,000 distinct AI agents that used 130 billion tokens over 88 hours (Dwarkesh Podcast). In the surrounding coverage, the problem is identified as the Navier–Stokes equations, one of mathematics' six Millennium Prize Problems, which describe how fluids move (The Neuron, 18 September 2026). Another summary characterizes the claim as a company claim and qualifies the agent count as a "purported" 10,000 (BigGo).
It is worth pausing on what this phrasing actually is: the host's rendering of OpenAI's announcement, not independently verified evidence. No primary OpenAI documentation — no paper, preprint, or official press release — and no confirmation from the Clay Mathematics Institute appears in the available source material. That does not mean the claim is wrong — but that "solved" is for now a framing from the podcast and newsletters, not an established mathematical fact. It is a difference that matters when the result is simultaneously being used to draw sweeping conclusions about where AI development stands.
Brown himself is a key source for interpreting it: according to the podcast, he was one of the key contributors behind what became o1 and OpenAI's reasoning models, and now works on multi-agent systems. The episode, together with The Neuron's follow-up of 18 September, provides the first thorough walkthrough of how the system works — and of Brown's own caveats.
The mechanism: why multiple agents at all?
Brown's argument starts from an empirical pattern, which he put this way in the transcript: "The way I think about it: when you plot the performance of these reasoning models with test-time compute on the x-axis and performance on pretty much any reasoning benchmark on the y-axis, you see a very clear pattern: the longer these models take to think about their answer, the better they do" (Dwarkesh Podcast). The quote comes from the Dwarkesh transcript and is reproduced here in English.
The problem is that this extra thinking time takes wall-clock time. A model given ten times more reasoning time gets better — but the answer arrives ten times later. According to the episode, Brown uses the analogy of a student with five minutes on the SAT versus five hours: given ample time, most people solve far harder problems, but in practice you have to split up the work and parallelize. Multi-agent systems are the AI version of this: instead of one agent reasoning serially for hours, you let many agents work simultaneously and exchange results.
So coordination is not the goal in itself — it is a workaround for a time bottleneck.
The architecture: minimal structure, coordination that emerges on its own
According to a summary of the episode (BigGo), OpenAI's approach departs from the traditional multi-agent architecture with rigid coordinator and worker roles. Instead, as little structure as possible is built in: each agent gets a primitive tool for sending messages to any other agent, and the message is inserted into the recipient's context. The coordination — who shares what with whom — purportedly emerges on its own.
This is the concrete technical news in the story. The point is that if you give the agents a simple communication primitive and a capable base model, the system can itself discover how it should organize the work — rather than humans predefining a hierarchy that may turn out to be wrong for each new type of task.
Note at the same time that this description comes from an automatically generated summary of the podcast, not from verbatim quotes in the transcript. It is probably credible, but the details should be clarified against the full transcript.
The numbers: sublinear scaling, not free parallelization
What do you actually buy by adding agents? The Neuron (18 September 2026), citing OpenAI's published scaling figures, reports the following: four agents working together complete a task roughly twice as fast, but at roughly twice the compute cost. Sixteen agents show similar, though somewhat less efficient, behavior.
In other words: the time gain is roughly linear in this range, but the cost scales too — and the gain flattens out as the swarm grows. This is sublinear economics, not free parallel acceleration. You buy time savings with tokens, and the margin per additional agent shrinks.
The second caveat Brown himself raises is that parallelizability is domain-dependent. Math problems with a natural decomposition may parallelize reasonably well; other task types where the sub-results are tightly coupled will likely benefit far less from the same trick. One data point from a Millennium Prize problem says little about how well the approach generalizes.
The caveats: one data point, no ablations
The most striking thing in the episode, according to the BigGo summary, is Brown's own credit distribution: the Millennium result should not be attributed primarily to multi-agent coordination — he puts less than 10 percent of the credit there — but to a genuinely powerful general model capable of working over long time horizons.
In other words: what made the run possible was first and foremost that the base model was good enough to hold the thread over 88 hours and 130 billion tokens. The swarm was about time, not about a new kind of intelligence.
Brown also points to a methodological gap: there is no ablation showing how long a single agent would have taken on Navier–Stokes, and no measurement isolating the gain of 10,000 agents versus 1,000 (BigGo). Without that kind of control set, you cannot say for certain how much the parallelization contributed, or whether scaling from 1,000 to 10,000 agents was worth the token bill. The result is — as Brown himself, per the summary, enumerates — one data point without the comparisons needed to generalize.
This is unusual humility from the researcher behind a big result, and it is the analysis's most important yield: the headlines read the announcement as a multi-agent breakthrough; the central figure himself reads it as evidence that long reasoning horizons in a strong base model are what really carries the result.
What remains unresolved
The biggest open question is the verification status. A claim that a system has "solved" a Millennium Prize problem is in principle something mathematicians can check. So far no such confirmation exists in the available material, and it is unclear whether the mathematical community has even reviewed the argument. Until an independent review exists, the careful phrasing is that OpenAI claims to have solved the problem, not that it is solved.
The second is methodological: which ablations would isolate the multi-agent contribution? At least three are relevant — a single agent on the same problem, a swarm of 1,000 agents, and one of 10,000 — and together they would yield an actual scaling curve. Such experiments are expensive in tokens, but they are the only way to move the discussion from anecdote to measurement.
Third, the episode's own chapters point beyond the result itself. Among the timestamped topics we find "what math progress tells us about recursive self-improvement" (Dwarkesh Podcast), along with discussion of whether models' chains of thought become harder to monitor as they degrade, and how one can know whether models are aligned before recursive self-improvement. The episode material thus suggests that mathematical results of this kind are treated as a data foundation for thinking about what happens when AI starts automating AI research — and that the monitorability of the reasoning is an open problem in that scenario.
Why this matters
The announcement has triggered a broader debate about lab safety and the pace of AI progress. That coverage is largely secondary, much of it cannot be verified in the source base, and should therefore be treated with caution.
The solid yield of the week is more down-to-earth. First: the scaling economics of multi-agent systems are bought, not found — four agents give roughly double the speed at double the cost, and the gain flattens out from there. Second: the minimal-structure architecture with a simple message primitive is a concrete, testable design decision that others can copy or challenge. Third: the most important contribution to the result was, according to Brown himself, probably not the swarm, but a base model that tolerates long horizons — a reminder that the most significant advances in reasoning AI are still happening in the model itself, not in the orchestration around it.
That is perhaps the most useful lesson from the 88 hours and the 130 billion tokens: before concluding that swarms are the future, you need the numbers showing what the swarm actually added.

