Stanford: The gap between Chinese and American AI models has shrunk from 1,300 to 39 points

Stanford's 2026 AI Index reported in October that the performance gap between the best American and Chinese models has narrowed to 2.7 percent. At the same time, evaluations, prices, and usage data show that leadership in reasoning…

Two almost equal-height towers of matte concrete blocks, one in cool gray and one in muted red, stand side by side with only a single thin block remaining between them – a visual metaphor for the shrinking performance gap between American and Chinese AI models.
Gift article

Stanford: The gap between Chinese and American AI models has shrunk from 1,300 to 39 points

Stanford's 2026 AI Index reported in October that the performance gap between the best American and Chinese models has narrowed to 2.7 percent. At the same time, evaluations, prices, and usage data show that leadership in reasoning, compute, and talent remains measurable — and uneven.

Autumn 2026 has delivered a dense cluster of data points about the US–China AI competition, and they point in directions that do not quite reconcile. According to Stanford's 2026 AI Index, as reported by Bloomberg and summarized by Crypto Briefing on October 4, the performance gap between the top models on each side was down to 2.7 percent, or 39 points, as of March 2026. In May 2023, that gap stood at more than 1,300 points. Much of the work of closing it, Bloomberg writes, was done by DeepSeek.

But 2.7 percent measures one thing: general benchmark performance. The same week the figure became known, other signals emerged that the other "gaps" — reasoning capability, compute, prices, talent, actual usage — are moving at different speeds, and in some cases not narrowing at all. This article separates what each number actually measures from what it is easily read as.

What the Stanford Figure Captures — and What It Doesn't

The AI Index figure (2.7 percent / 39 points as of March 2026) is an aggregate of benchmark performance. Worth noting: the available reporting does not specify which benchmark aggregate the figure is built on, or which models are compared. That means it should be read as an expression of top models now scoring nearly identically on general tests — not as proof of overall parity.

The timeline behind the figures is nevertheless concrete. DeepSeek launched its reasoning model R1 in January 2025, a release that per Crypto Briefing put Chinese AI visibly on the global map. The V4 series began rolling out in April 2026, with open weights, and V4-Pro is priced at roughly $3.96 per million output tokens.

Where the US Still Holds a Measurable Lead

A May 2026 evaluation from CAISI found that even the best Chinese models lag 6–8 months behind American frontier models on complex reasoning and cyber tasks, as relayed by Crypto Briefing. DeepSeek V4 Pro was, according to CAISI, roughly eight months behind. So the question is not whether Chinese models will get there — it is how long the delay is on the heaviest tasks. The same caveat applies: the evaluation is known here only through secondary reporting.

On infrastructure: the US accounts for roughly 74 percent of compute and leads on high-impact patents. China leads on publication volume and some citation metrics. These are different kinds of leadership, but they are not equivalent: publication volume is easier to scale than access to computing capacity.

Usage and Price: The Numbers That Complicate the Picture

OpenRouter data (via Crypto Briefing) show that Chinese models accounted for 50–67 percent of token processing on the platform in mid-2026. OpenRouter routes developer requests across many models, so the figure tempts toward one conclusion: developers prefer Chinese models. But an independent analysis of the same platform (published October 3, 2026) questions that conclusion. The analysis found that the higher a model's roleplay share, the smaller its usage drop on weekends — "it is completely monotonic," the author writes. deepseek-v3.2, with an 87 percent roleplay share, had a weekend drop of only 0.966. That suggests a substantial portion of Chinese model usage on OpenRouter is personal/entertainment use with high loyalty — not necessarily workloads where developers actively choose Chinese models for technical reasons. The analysis covers only OpenRouter usage, not the entire LLM market.

On price: the Jefferies analysis (reported by the Wall Street Journal on October 1, relayed via MSN/WSJ) found that the average price gap between American and Chinese models widened from 60 percent in August to 70 percent in September 2026 — Chinese models average 30 percent of the American price. But that is a tendency, not a rule: Jefferies simultaneously points out that GPT-6 Luna is 74 percent cheaper than DeepSeek V4.1 Flash. An American model cheaper than a Chinese one breaks with the "Chinese = cheap" narrative and shows the pricing market cannot be read across tiers with a single rule.

Talent: The Reversal Worth Taking Seriously

A study from Carnegie China, covered by the New York Post on September 26, 2026, claims that China hired 41 percent of the world's leading AI researchers in 2025, versus 34 percent who went to American companies. That is a reversal from 2022, when the figures were 46 (US) versus 27 (China). The study also says 57 percent of elite researchers completed their undergraduate education in China.

Worth noting: the Post frames the figures with "reportedly," and the study itself is not in the source material for this article — what we have is the news coverage of it. South China Morning Post, cited by the Post, points to a concrete mechanism: DeepSeek and MoonShotAI now offer salaries that can compete with Silicon Valley. That explains how China can attract back researchers trained in the country without pay levels being an argument against it.

One Usage Example, Clearly Labeled as Anecdote

Takahiro Anno's first-person report from 26 interviews in Paris in September 2026, published September 30, describes the French government's cross-departmental AI gateway: it does not provide access to closed American models (Astra, Fable, Opus, Gemini), and instead runs open models on France-certified cloud infrastructure. DeepSeek is reportedly the most-used model, apparently restricted to coding tasks. This is an anecdote — Anno's own text is cautious ("it appears"), the text is machine-translated, and it covers only coding use. It is not confirmation of any broader pattern, but it illustrates one concrete way open Chinese weights are actually being taken into use in the public sector.

The Open Questions

Several things remain before the various numbers can be assembled into one picture:

  • Which benchmark underlies the 2.7 percent? Without that specification, the figure cannot be directly compared with CAISI's 6–8 month lag, which measures something else (complex reasoning, cyber tasks).
  • OpenRouter methodology. The 50–67 percent range is wide, and the underlying basis is unclear. Independent analysis suggests a large share of usage is roleplay — which is real usage, but a different kind than "developer adoption."
  • The pricing data. Jefferies' 70 percent average gap and DeepSeek V4-Pro at $3.96 per million tokens are not obvious to reconcile, and the GPT-6 Luna example shows prices vary sharply across model tiers.
  • Whether the CAISI lag holds. The evaluation is from May 2026, before any of the V4 series' later releases may have changed the picture. Whether 6–8 months is a stable gap measure or a snapshot is an open question.

The practical conclusion is that "the gap is narrowing" is true but incomplete: frontier parity on general benchmarks is real, while leadership on reasoning-heavy tasks, compute, and top-tier talent retention remains measurable and unevenly distributed.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Get the best of AI MAG in your inbox

News, analysis, and ideas at the intersection of AI and society.