Persistent Kernels: How Software Can Give GPUs Near-Cerebras Speed

OpenAI's Ultrafast mode has reportedly moved from Cerebras' wafer-scale silicon back to NVIDIA GPUs — and that raises a technical question that is less about raw compute than about how much data must be moved for each generated token.

Illustration: a continuous dark silicon block fused onto a metal plate, surrounded by short tightly coiled copper threads — an image of computation staying resident instead of moving data.
Illustration
Gift article

Persistent Kernels: How Software Can Give GPUs Near-Cerebras Speed

OpenAI's Ultrafast mode has reportedly moved from Cerebras' wafer-scale silicon back to NVIDIA GPUs — and that raises a technical question that is less about raw compute than about how much data must be moved for each generated token. Here is the mechanism behind low latency, and why software may catch up with what others need dedicated hardware to achieve.

The news event: Ultrafast switches hardware

In August 2026, OpenAI announced an Ultrafast mode for GPT-5.6 Sol — according to an article on the Japanese platform note.com, which is the sole source behind this story. OpenAI itself described the mode as "Powered by Cerebras," the author writes, adding that this irritated NVIDIA shareholders. The speed reached up to 750 tokens per second in output, up to 14 times faster than the standard version.

Now the picture reportedly looks different: the semiconductor analysis firm SemiAnalysis has reportedly pointed out that GPT-6.1 Sol Ultrafast "runs on NVIDIA GPUs at low batch sizes, not Cerebras." This is a competitively charged claim from an analysis firm, reported through the same secondary source, and it has not been independently verified. But if it holds, it points to a concrete technical question: How can ordinary GPUs even come close to wafer-scale hardware when it comes to latency per token?

The answer, as SemiAnalysis reportedly sketches it, lies in two parts: an explanation of why GPUs underperform at batch size 1, and a verification of software that removes much of that overhead.

Why every token is a trip to memory

An autoregressive language model generates one piece of text — one token — at a time, and each token depends on all the ones before it. As the note.com article explains it (translated from English): "Every time the next word is generated, it must access a large amount of data such as model weights and the KV cache, perform computations, and then produce the next word."

The point is that the actual computation per token is relatively small. The large model weights must, however, be read from memory for every single token, along with the KV cache — the intermediate store of previous token representations. When a GPU generates thousands of tokens in parallel (high batch size), the same weight movement can be reused for many computations, and the GPU's enormous compute is fully utilized. But at batch size 1 — one user waiting on one answer — it is memory movement that sets the pace, not compute. Decoding is then "memory-bandwidth-bound": speed is limited by how many terabytes of data can be fetched per second, not by how many trillion operations the chip can perform.

That is why the debate over fast AI inference is, at bottom, a debate about memory bandwidth.

The GPU's weakness at batch size 1

If memory bandwidth is the bottleneck, then modern GPUs should in principle do well. NVIDIA's Blackwell Ultra generation (B300/GB300) delivers, according to figures cited in the source, roughly 8 TB/s of HBM3E bandwidth per GPU. An entire GB300 NVL72 rack reaches up to 576 TB/s of HBM bandwidth across 72 GPUs, and the NVLink interconnect reaches 130 TB/s per rack.

The problem, according to SemiAnalysis as relayed in the article, is that the GPUs never get to use that bandwidth in ultra-low-latency scenarios. "While GPUs are extremely strong at high throughput, they cannot fully exploit their inherent memory bandwidth in ultra-low latency inference such as batch size 1," SemiAnalysis analyzes, according to the article. The causes are listed as small delays tied to four things:

  • GPU kernel launches: every small operation in the network must be started individually.
  • Terminations of those same kernels.
  • Synchronization between operations.
  • Writing intermediate data back to HBM — the GPU's external high-speed memory — between each operation.

In normal GPU inference, a series of small processes runs sequentially, across thousands of GPU kernels. Every launch, every termination, and every round of intermediate storage costs microscopic delays. At high batch size, these costs are amortized over enormous amounts of work. At batch size 1, they make up a substantial share of the time between each token — and the GPU effectively sits waiting instead of reading weights.

The hardware alternative: Cerebras and wafer-scale

Cerebras took a different path. The company's WSE-3 uses, according to the company's specifications as cited in the note.com article, nearly an entire silicon wafer as one massive processor, with roughly 44 GB of SRAM placed directly next to the compute cores. The on-chip memory bandwidth is officially stated at around 21 PB/s.

The comparison is striking: 21 PB/s on-chip versus roughly 8 TB/s of HBM per B300 GPU — a factor of over a thousand. The point is not just raw bandwidth but the architecture. Because the model's weights can be held in SRAM beside the compute cores, Cerebras avoids much of the constant shuttling of data back and forth between processor and external memory that the GPU architecture requires per token. The wafer-scale design essentially builds the bottleneck out of the system.

It was against this picture that OpenAI's reported "Powered by Cerebras" label and 750 tokens per second were read as a direct exploitation of the architectural advantage at low batch size.

The software answer: TileRT and persistent kernels

But the architectural advantage is not the whole story. In August 2026, SemiAnalysis verified, according to the note.com article, low-latency inference software called "TileRT" on an NVIDIA B200. The design is simple to describe: instead of launching thousands of sequential GPU kernels, TileRT keeps the model's entire decoding process resident on the GPU as one massive "persistent kernel."

Mechanically, this means the code running the decoding starts once and stays on the chip for the whole generation. The intermediate steps that would normally each require their own kernel launch, their own termination, their own synchronization, and each write of intermediate data to HBM instead happen inside the same, persistent process. The full list of overhead costs that SemiAnalysis identifies at batch size 1 is thus addressed by software — without changing the hardware.

This is the core of the argument, as the source presents it: the speed gap between GPUs and wafer-scale hardware at low batch size is partly a software problem, not only a hardware problem. The GPU does not need 21 PB/s to approach Cerebras-class behavior; it just has to stop wasting its time on its own administration between operations.

What we don't know — and what is analysis

It is worth being honest about the evidence base. All the claims in this story — the OpenAI announcement, the SemiAnalysis analysis, the TileRT verification, and the claim that GPT-6.1 Sol Ultrafast runs on NVIDIA — stem from a single Japanese note.com article. Neither SemiAnalysis' report, OpenAI's statements, nor TileRT documentation is available here in primary form, and none of the 2026 events has been independently confirmed.

For the reader's clarity: what in this article is labeled as SemiAnalysis' analysis, OpenAI's claims, or company specifications is relayed as the secondary source reports them. The explanatory passages on memory-bandwidth-bound decoding, the amortization of overhead at high batch size, and the significance of the SRAM architecture are AIMag's own attempt to make the source's — and SemiAnalysis' purported — reasoning understandable; they rest on the same single source and have not been independently verified. The conclusion at the end of the article is editorial analysis of what the source reports, not documented fact about NVIDIA's product strategy.

One note on the source's scope: the available material covers persistent kernels, low batch size, and memory bandwidth. The source's further analysis of what token speeds NVIDIA actually achieves with such techniques is cut off in the source text and cannot be relayed here.

There are also open technical questions: how large a share of GPU overhead persistent kernels actually remove in practice, what trade-offs TileRT makes in flexibility (a fixed, monolithic kernel is less flexible than a kernel sequence that can be changed), and how this software approach scales to larger models and batch sizes above 1 — where GPUs, according to SemiAnalysis, are already "extremely strong."

The big picture is nevertheless clear enough to understand the debate, as the source sketches it: the speed of fast AI answers is determined to a small degree by how many computations a chip can perform, and to a large degree by how quickly data can be fetched — and by how much time the software manages to save between fetches. Cerebras solved it in silicon. According to the reporting this story rests on, NVIDIA is in the process of answering in software.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Get the best of AI MAG in your inbox

News, analysis, and ideas at the intersection of AI and society.