China or the US ahead? LiveBench and CAISI give conflicting answers on AI catch-up

A Bloomberg Intelligence note claims the gap between China's and the US's best models has narrowed to 3 percent, driven by DeepSeek V4.1 Flash. But NIST's evaluation unit CAISI concluded in May that the difference was eight months.

Two nearly equal stacks of gray gauge blocks with a caliper wedged between them, illustrating a small but disputed measurement gap in the AI race between China and the US.
Illustration
Gift article

China or the US ahead? LiveBench and CAISI give conflicting answers on AI catch-up

A Bloomberg Intelligence note claims the gap between China's and the US's best models has narrowed to 3 percent, driven by DeepSeek V4.1 Flash. But NIST's evaluation unit CAISI concluded in May that the difference was eight months. Both could be right — and that is precisely what makes the measurement landscape so hard for businesses to navigate.

In early October 2026, Robert Lea, senior analyst at Bloomberg Intelligence, published a note documenting a record-low 3 percent gap between China's and the US's best AI models. The driving force is DeepSeek V4.1 Flash, released September 10. According to the LiveBench snapshot from October 4, DeepSeek scores 77.3 on agentic coding, against Anthropic's Claude Fable 5.1 Max Effort at 66.1. On the overall LiveBench score, it stands at 81.1 versus 83.4 — a difference of 2.3 points, which rounds to roughly 3 percent of Anthropic's score.

The number is worth taking seriously, but also worth placing precisely. It comes from an analyst's note, relayed through secondary sources, and it describes a gap measured at one specific benchmark snapshot. NIST's Center for AI Standards and Innovation (CAISI), in its May evaluation of DeepSeek V4 Pro — the predecessor to V4.1 Flash — concluded with something entirely different: roughly an eight-month development lead for the US, measured across nine pre-selected benchmarks spanning five domains.

This is not a contradiction that can be resolved editorially. It is a genuine methodological disagreement, and it says something important about how little "the gap between China and the US" is a single text that can be read unambiguously.

What the numbers actually show

The gap history Lea traces is itself news. Earlier in 2026, the gap stood at around 15 percent. In May, it was down to around 9 percent. Now, after V4.1 Flash, it is 3 percent.

Agentic coding is the dimension attracting the most attention, because it breaks with a persistent argument: that the West leads precisely where enterprise customers actually use the models — coding within agentic workflows. The LiveBench snapshot from October 4 shows DeepSeek 11.2 points ahead of Anthropic on this dimension, while the gap on overall score is minimal.

But there is a second layer of numbers that requires greater caution. DeepSeek's own technical report claims that V4.1 Flash scores 74.2 percent on DeepSWE v1.1, described as a demanding agentic coding benchmark. That places the model ahead of Anthropic's Opus 5 (74.0 percent) and OpenAI's GPT-5.6 Sol (73.0 percent). The reporting itself, however, flags that these figures are self-evaluation, and that no independent third party has published validation of them. These are two different epistemic statuses: the LiveBench figures are third-party measurements cited by an analyst; the DeepSWE figures are the manufacturer's own claims. They should not be conflated.

What V4.1 Flash is

Part of the explanation for why this particular model moves the numbers lies in technical choices that cut the costs of inference.

V4.1 Flash uses a Mixture of Experts architecture with 748 billion parameters in total — 552 billion in the core backbone itself and 196 billion Engram parameters, a new component for persistent memory introduced in this model generation. During inference, only a sparse selection of the parameters is activated per token, which keeps computing costs down despite the large total model size.

The model supports a context window of up to 1 million tokens with output of up to 384,000 tokens, and adds native image understanding that its predecessor lacked. Pricing is listed at $0.30 per million input tokens and $1.20 per million output tokens at peak load, with half price outside peak load.

The combination of a large context window, sparse activation and aggressive pricing is classically aimed at exactly the agentic coding workflows where LiveBench now shows leadership. Agentic flows generate many tokens and require long contexts — and there, price matching and architectural choices feed directly into benchmark performance and usage cost.

Why the two measurements conflict

The CAISI evaluation from May and Lea's LiveBench analysis differ on three points, each of which is enough to explain the divergence.

First: different models. CAISI evaluated V4 Pro, not V4.1 Flash. Since V4.1 Flash was released after the CAISI evaluation, the two analyses measure different points in the technology's development.

Second: different times. May is not October. In an industry where the gap history, according to Lea, moved from 15 to 9 to 3 percent in half a year, an eight-month-old estimate is by definition a historical document.

Third: different methods. CAISI used nine pre-selected benchmarks across five domains — a broader, more locked-down way of reaching a conclusion. LiveBench provides a snapshot of one leaderboard composition. One method attempts to capture general development distance; the other captures current peak performance on a specific ranking. Both can be methodologically defensible and still arrive at different answers.

That means the question "has China caught up with the US?" does not have a single answer to retrieve. It has at least two, and both are worth knowing.

Breadth and the bottom

A nuance easily drowned out in the headlines: the catch-up is happening at the frontier, not across the field. Only three of the top 15 models on LiveBench are Chinese. Lea also points out that the Chinese market now hosts over 1,100 large language models, and that he does not expect the Chinese AI sector to become profitable before 2030 at the earliest.

In other words: a 3 percent gap at the top coexists with fragmentation at the bottom and an unresolved profitability question. Lea himself attributes the catch-up to Chinese labs' technical progress and optimization for domestic hardware — which he argues raises questions about how effective US chip export controls actually are. That is a consequential inference, not a measured result, and it should be read as such.

What an enterprise buyer can and cannot conclude from this

The benchmark landscape itself offers buyers three concrete lessons.

One: distinguish self-evaluation from independent measurement. The DeepSWE figures from DeepSeek are factory-reported. The LiveBench figures are third-party measured. A purchase based on agentic coding should demand documentation from an independent measurement, not a technical report from the seller.

Two: distinguish frontier from field. Three Chinese models among the top 15 is not the same as the Chinese model lineup generally being on par with the American one. The catch-up is real, but narrow.

Three: distinguish benchmark from production behavior. Agentic coding on a leaderboard measures short, defined tasks under controlled conditions. Production use involves reliability, security, integration and cost over time — conditions that none of the cited measurements capture.

Open questions

The honest accounting: everything in this story rests on secondary sources. Lea's note from Bloomberg Intelligence, the underlying Bloomberg reporting, the LiveBench snapshot from October 4 and DeepSeek's technical report are not available in the source material here — they are cited through three secondary coverages, two of which are nearly identical syndications of the same article. That does not make the figures wrong, but it means none of the key figures can be checked against the primary source in this story.

It also means the open questions are concrete: Will an independent third party validate the 74.2 percent DeepSWE v1.1 figures? Does LiveBench's composition capture what enterprise buyers actually care about, or does CAISI's broader domain measurement give a truer picture of the distance? And if Lea is right that optimization for domestic hardware is driving the catch-up, what does that mean for the assumptions behind the export control regime — and for how quickly the gap could shift again?

What can be stated with reasonable confidence is this: on one third-party measurement, on one date, the best Chinese offering is very close to the best American one — and ahead on agentic coding. How much that matters for development distance, for vendor selection and for policy depends on which of the two measurements one weights most heavily. And that is exactly what neither source can decide for you.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Get the best of AI MAG in your inbox

News, analysis, and ideas at the intersection of AI and society.