709 vs 452 tokens per second: Inferact's own measurements place TPU v7 ahead of GB200 on Kimi K3

The company Inferact, which counts several vLLM maintainers on its team, has released open "megakernels" for Google's TPU v7, also known as Ironwood.

Illustration: Two industrial processor towers side by side, the taller, warmly lit one representing higher throughput in a chip performance comparison.
Illustration
Gift article

709 vs 452 tokens per second: Inferact's own measurements place TPU v7 ahead of GB200 on Kimi K3

The company Inferact, which counts several vLLM maintainers on its team, has released open "megakernels" for Google's TPU v7, also known as Ironwood. According to the company's own measurements, relayed by SemiAnalysis and reported by OfficeChai on September 28, 2026, 16 TPU v7 chips deliver 709 tokens per second per user on Moonshot AI's Kimi K3 — versus 452 tokens per second for 16 Nvidia GB200 GPUs running vLLM's published recipe. That amounts to roughly 1.6 times the throughput, or 56% better, as SemiAnalysis put it.

The numbers come with significant caveats: They originate from Inferact itself, which wrote both the kernels and the comparison methodology. The headline figure depends on speculative decoding, the GB200 baseline is vLLM's standard recipe rather than a hand-optimized kernel, and the results apply only to low batch and a single user. But the underlying software and hardware story — how a TPU program can keep an entire model layer's weights in on-chip memory — may be the most important news, especially now that SemiAnalysis describes Google's work to make TPUs usable outside its own data centers as "full steam ahead".

Megakernels in Pallas

The megakernels are written in Google's programming language Pallas and pack all 92 of Kimi K3's MoE layers (mixture-of-experts) into a single call distributed across 16 chips. It is a radical approach: Instead of chaining each layer as a series of smaller operations, with intermediate synchronization and memory traffic, the entire decode phase is directed as one program across the whole chip array.

This is where the hardware differentiation comes in. Each TensorCore has 64 MiB of software-managed scratchpad memory called VMEM — according to Inferact's stated specifications, large enough to place most of a layer's weights in advance, so that computation doesn't have to wait for fetches from slow main memory. An entire GB200 GPU, by comparison, has around 38 MiB of comparable tensor memory in total, spread across 152 streaming multiprocessors.

Beyond memory, the chips are nearly even. Inferact lists 2.31 petaflops of BF16 compute and 7.38 TB/s of memory bandwidth for the TPU, versus 2.5 petaflops and 8 TB/s for the Nvidia chip. With a slim margin to Nvidia on raw specifications, the difference in decode speed points to one place: how much of the working set can live near the compute units, and how much of that placement the software itself can control.

The Qwen result and the Cerebras question

On Alibaba's Qwen 3.8 27B, four TPU v7 chips reached 1,515 tokens per second per user, versus 695 on four GB200s, according to Inferact's data. SemiAnalysis' chart placed Cerebras at around 1,850 tokens per second per user on the same model, which led the analysis firm to say that four TPUs had come "close" to a single giant Cerebras wafer on interactivity.

This is the story's weakest claim. The Cerebras figure in SemiAnalysis' chart is marked as approximate, with no accompanying methodology, and Cerebras has not responded. Concluding that TPUs are "closing in" on Cerebras thus rests on an estimate without a documented setup, measured against one company's own benchmark. It's a difference the reader should carry with them.

How much of the lead is speculative decoding?

The headline figure of 709 tokens per second rests on speculative decoding with the DSpark method, at an acceptance length of six — the technique guesses several tokens per step, and accepted guesses count as multiple generated tokens. At an acceptance length of three, the gap was smaller: 350 versus 229 tokens per second. Each decode step took around 8.5 milliseconds. Without speculative decoding, the kernels deliver, according to the reported figures, roughly 1.4 to 2 times GB200's decode throughput at batch sizes one to eight.

That means the real, technique-independent lead is substantial — but far more modest than 56%. The headline number assumes both that speculative decoding with a high acceptance length works well on the workload, and that the competition is not optimized.

And the competition is not optimized. The GB200 baseline is vLLM's standard recipe on 16 GPUs — not a specially written kernel, and not a full NVL72 rack with Nvidia's fastest interconnect topology. A fair comparison would pit two parties' best software against each other; here a bespoke megakernel meets a generic execution schedule.

What the results actually measure

The results apply to low batch and a single user — that is, interactive speed, which is where latency matters most for agents and coding tools. They say nothing about maximum throughput with many concurrent users, which is the economically most important metric for most inference services. The kernels' advantage at batch 1–8 is documented in the reported material; at high batch sizes no comparable data exists.

Accuracy preservation is also documented only by Inferact's own figures: 0.944 on GPQA-Diamond and 0.972 on GSM8K. No independent verification of these numbers is available in the accessible material.

Why it's happening now

Inferact's release coincides with SemiAnalysis' broader assessment that Google's "externalization of software" around the TPUs is "full steam ahead" — that Google is actively working to make the chips usable outside its own data centers. The analysis firm had previously found that Ironwood delivers up to 50% better performance per dollar than Nvidia's Blackwell Ultra, and TPU adoption has grown among labs looking for alternatives to Nvidia's chips.

In that context, the open megakernels are more than a benchmark: They show that the TPU software stack is mature enough for third parties to write performance-critical code in Google's Pallas and beat reference implementations on the Nvidia side. For buyers weighing TPUs against Nvidia alternatives, such third-party benchmarks are useful — but precisely for that reason, they deserve thorough scrutiny.

Open questions

None of the numbers in this story are independently verified. The benchmark was authored by the vendor, which has an obvious interest in the results and wrote both the kernels and the comparison methodology. There is no primary documentation — Inferact's release, SemiAnalysis' dashboard, or the source code repository — that has been reviewed in this evidence base; all figures rest on one secondary article relaying the company's own results (OfficeChai).

The most important open questions: How do the kernels perform at high batch sizes and real multi-user traffic? What would an equivalently optimized GB200 kernel deliver? And does the Cerebras comparison hold up under closer examination? Until such tests exist, the safest thing one can say is that Inferact has shown that TPU v7's architecture — with 64 MiB of software-controllable memory per TensorCore and a programming language that exploits it — can win on single-user decoding under the right conditions. The rest remains unanswered.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Get the best of AI MAG in your inbox

News, analysis, and ideas at the intersection of AI and society.