Back
AI News

Paying for frontier AI now buys a 4.4-month head start — at five times the cost, Mozilla report finds

Mozilla's new report on open source AI finds that the best Chinese open models now trail US closed frontier models by only 4.4 months. The advantage is not universal: closed models earn their price especially on expert work, demanding…

AIMag.no
AIMag.no
September 16, 2026 · 4 min
Two nearly identical paper blocks, one tall and one slightly shorter, with a red measuring ruler spanning the small gap between them — an illustration of Chinese open models nearly catching up with American frontier models at a fraction of the price.

Paying for frontier AI now buys a 4.4-month head start — at five times the cost, Mozilla report finds

Mozilla's new report on open source AI finds that the best Chinese open models now trail US closed frontier models by only 4.4 months. The advantage is not universal: closed models earn their price especially on expert work, demanding retrieval, and long context — and the divide falls roughly between tasks of eight and twelve hours.

Mozilla published its State of Open Source AI report on September 15, 2026, and Ars Technica has had an exclusive preview. All the figures below come from Ars' preview; AIMag has not seen the report itself.

The number that matters: 4.4 months

The report's headline finding is that the performance gap between US closed frontier models and the best Chinese open-weights models has shrunk to 4.4 months. For an industry where models improve rapidly, that is a short distance — short enough that the question is no longer whether open models reach the frontier, but when, and what you are actually paying for in the meantime.

The concrete example is Moonshot AI's Kimi K3. According to the report, the model sits just three points behind Anthropic's closed frontier model Fable 5 on the Artificial Analysis Intelligence Index — at roughly 30 percent of the price. That is what underlies the math of five times the cost for a relatively small performance gain.

Where the divide actually falls

The gap is not uniform, however. The report's most practical figure concerns task length: the best closed model can reliably complete a job 1.7 times as long as the longest job the best open model can manage. The measurements build on METR's time-horizon metric, which measures how long a human would need for a task the model can complete — a metric that has doubled on a cadence that has accelerated over time (as described by Ars, with reference to METR).

The distribution, according to Mozilla as relayed by Ars, looks roughly like this: the best open models reliably complete tasks of up to around eight hours. In the span between eight and twelve hours, the difference becomes decisive — that is where closed models still manage jobs that open models do not reliably complete.

Mozilla's chief technology officer Raffi Krikorian frames it as a division-of-labor question: "[A closed model] earns its premium in a few places: expert professional work, high-intensity retrieval, and long context," he said in an email to Ars.

That also explains why organizations still pay. Closed models work out of the box and come with what Krikorian calls "compliance packaging, support, and accountability" — and many organizations lack the staff required to run open-weights models well.

How a real user splits the work

The report's illustration is DoorDash. According to Ars, citing Mozilla's report, the delivery company uses Kimi for routine work, while Fable is reserved for harder tasks that would normally take human experts longer. It is one cited example from the report — not a documented industry pattern, and neither Mozilla's figures nor DoorDash's usage has been independently verified by AIMag.

The pattern is still worth reading: the same company uses both model types, allocated by task, not by ideology.

The caveats

All the figures come via a single secondary source — Ars Technica's preview of a report AIMag has not seen. Ars also flags that custom test setups built by AI labs can produce inflated benchmark scores compared with third-party setups. Ars' text is also truncated where the comparison methodology is introduced, so the full picture of how the various measurements are equalized is unknown.

With that caveat, the report's message is clear: the premium for closed frontier models is real, but narrow — it is measured in months, not years, and it applies not to all tasks, only the hardest ones.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Sources

  1. Exclusive: Paying for frontier AI models buys 4-month head start at 5x the cost - Ars Technicaarstechnica.com