Ai2 launches Olmo-core 3: Open MoE framework reports 2.7 times higher throughput

Ai2 has launched Olmo-core 3, an open-source framework for training large mixture-of-experts models. The company's own measurements show roughly 2.7 times higher training throughput and configurations beyond a trillion parameters — but…

Illustration: dozens of thin copper wires routed through a small lattice of brass gates, only a few exiting densely packed — a metaphor for mixture-of-experts routing and high training throughput.
Illustration
Gift article

Ai2 launches Olmo-core 3: Open MoE framework reports 2.7 times higher throughput

Ai2 has launched Olmo-core 3, an open-source framework for training large mixture-of-experts models. The company's own measurements show roughly 2.7 times higher training throughput and configurations beyond a trillion parameters — but with one important caveat: the numbers measure system capacity, not the quality of a trained model.

The Seattle-based research organization Allen Institute for AI (Ai2) announced the framework on Thursday, October 1–2, 2026, and the project is already available on GitHub for developers and the open-source community, according to SiliconANGLE. Olmo-core 3 is the training infrastructure behind Ai2's next-generation OLMo models, and the release means independent labs can now evaluate the approach for themselves.

What actually changed: from FSDP to DDP

The core of the update is a choice of parallelization strategy. Earlier versions used fully sharded data parallelism (FSDP), in which model weights are split invisibly across GPUs and gathered back together when needed. That gives good memory utilization, but the cost is repeated weight-gathering operations during training.

Olmo-core 3 switches to a stack based on distributed data parallelism (DDP) in which the experts are kept resident on the GPUs, and the relevant data is routed to them instead, as Unite.AI writes. Since the weights are not gathered anew every time they are needed, one of the major communication costs of MoE training is eliminated. In a MoE model, a router selects only a few expert networks out of many for each token, and it is precisely this routing that makes it possible to increase total parameter capacity without increasing the compute per token.

The numbers — all from Ai2 itself

Ai2's own benchmark figures make up the bulk of the launch documentation, and they should be read as the company's claims:

  • 52,000 tokens per second per GPU versus roughly 19,400. In a preliminary test on eight NVIDIA B300 GPUs, Ai2 reports that a 47-billion-parameter MoE model achieved 52,000 tokens/s per GPU on the new stack, versus 19,400 on the previous implementation — what Ai2 describes as roughly 2.7 times throughput (Unite.AI). One source discrepancy is worth noting here: SiliconANGLE frames the comparison as being against Nvidia's established Megatron-core, while Unite.AI sets it against Ai2's earlier implementation. Which baseline underlies the 2.7x figure is therefore not settled in the available sources.
  • Scaling the expert pool. In one benchmark, the number of experts grew from 8 to 128, still with four experts selected per token. That kept active parameters roughly fixed at 3.2 billion while total capacity rose from 4.6 to 47 billion — with a training throughput loss of under 5 percent, according to Ai2 via Unite.AI.
  • Trillion scale. Ai2 states that it benchmarked a configuration with 1.2 trillion total parameters and 58.36 billion active per token across 512 GPUs, with the highest observed throughput at 858 TFLOP/s per GPU (Unite.AI).
  • 2.38 trillion with DeepEP v2. Tests with DeepEP v2, an alternative communication layer for routing between experts, reached a configuration of roughly 2.38 trillion total parameters, TechTimes reports.
  • Lower precision as a lever. In a controlled benchmark on four B300 GPUs with even expert loading, enabling MXFP8 in the most favorable parts of the system yielded roughly 21 percent higher throughput than BF16, while peak active memory fell from around 103 GiB to 95 GiB (TechTimes).

One smaller detail is also unresolved: SiliconANGLE writes "Nvidia B3000" as the GPU model, while the two other sources write "NVIDIA B300." The exact model name cannot be determined from the available source material.

The caveat that comes with the numbers

Ai2 itself is clear about what the benchmark figures mean and do not mean. Both headline benchmarks used random routing — an artificial substitute for real token distribution that lets the system's throughput be measured without an actual model being trained. They demonstrate what the hardware stack can handle; they are not proof that Ai2 has trained a trillion-parameter model, or that such a model would perform well, TechTimes writes of the technical report.

That means "trillion-parameter scale" in this context is a systems reference, not a quality badge. The technical report should, according to TechTimes, be characterized such that the headline numbers are systems references under random routing, not claims about time-to-quality. For readers comparing such numbers across frameworks, the distinction is essential: throughput under synthetic routing says nothing about how a model converges or what it learns.

Why it matters

The big picture is that training infrastructure at this level is normally developed in-house at frontier companies and rarely shared. That an open stack reports system configurations of 1.2–2.38 trillion parameters, with documented throughput of up to 858 TFLOP/s per GPU across 512 GPUs, gives labs and developers without frontier companies' resources a concrete reference point — and code they can run and build on via GitHub.

The payoff of the DDP setup is also of more general interest: keeping expert weights fixed on the GPU and moving data instead is a design choice that could inspire other open training projects, particularly where FSDP's repeated weight gathers dominate the cost. The combination with MXFP8, which according to Ai2 delivered 21 percent throughput and lower peak memory use, points to lower numerical precision becoming an increasingly important part of training economics.

Open questions

Several things remain before the numbers can be taken as established fact:

  • Independent replication. All figures come from Ai2 itself, relayed through secondary sources. None of the outlets ran their own measurements.
  • The comparison baseline. Whether the 2.7x figure is measured against Nvidia's Megatron-core or Ai2's own previous implementation is unsettled, and whether the comparison was run on identical hardware and workload configurations cannot be verified.
  • Real routing. The systems references under random routing say nothing about performance or quality under real training loads. Today's synthetically routed benchmarks load the communication layers evenly; real routing patterns can create imbalances that change the picture.
  • The next step for OLMo. The framework is set to power Ai2's next generation of OLMo models. Those models — and what they actually achieve — will be the real test of whether the architectural overhaul also translates into quality.

For now, the release stands as a plausible but unverified advance in open MoE training: a DDP-based redesign with large self-reported numbers, immediately available for anyone who wants to put them to the test.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Get the best of AI MAG in your inbox

News, analysis, and ideas at the intersection of AI and society.