Quail, from Modal and CMU, claims a billion tokens per minute on one H100 GPU

Inference engine developer Modal and database researchers at Carnegie Mellon University have built Quail, an open-source inference engine that exploits known SQL structure to reorder requests for better KV caching.

Illustration: hundreds of wooden pegs scattered across a table, gathered into one tight ordered column in a groove, evoking reordered requests for better caching.
Illustration
Gift article

Quail, from Modal and CMU, claims a billion tokens per minute on one H100 GPU

Inference engine developer Modal and database researchers at Carnegie Mellon University have built Quail, an open-source inference engine that exploits known SQL structure to reorder requests for better KV caching. According to its own benchmark, it handles more than a billion tokens per minute on a single H100 GPU — over ten times faster than their vLLM baseline on one query. All figures, however, are self-reported and await independent verification.

The news

On 24 September 2026, Modal and Carnegie Mellon University's Full Stack Data Lab published the announcement of Quail — short for QUery-Aware Inference Layer. The release consists of three parts: the engine itself, a new benchmark for AI-SQL queries, and an installable package called quail-engine (version 0.1.0 on Modal) that lets readers replicate the demo themselves. The demo uses the qwen3-4b-fp8 model on a single H100 GPU.

The project is a collaboration between inference researchers at Modal and database researchers at CMU's Full Stack Data Lab. The blog post was written by Charles Frye and Shreya Shankar.

The problem: AI-SQL hits engines optimized for the wrong workload

The project's starting point is a workload type the authors call AI-SQL: LLM calls used inside large-scale database calls, such as classification or extraction over millions of rows. A single such query can consist of potentially millions of sequences of thousands of tokens each.

According to the authors, today's inference engines are optimized for a different usage pattern — agentic inference, where each request is long, unpredictable and, relatively speaking, independent of the others. When a massive AI-SQL workload is sent straight into such an engine, the result is poor caching and high overhead on the host process. Quail was built to solve exactly this: "So we built an inference engine to fix this: the QUery-Aware Inference Layer (Quail)," the authors write.

The mechanism: structure up front

The core technique rests on the query planner and the LLM inference layer talking to each other. When the planner knows the structure of a SQL query in advance — for example, that many rows will pass through the same template with shared prefixes — the inference engine can exploit that. The authors point to three mechanisms:

  • Query-aware ordering of requests. With a structured query in hand, the engine can order requests so that the KV cache is used better and eviction (what gets thrown out of the cache) happens in a planned way. According to the authors, this is "the big win": "with a structured query in hand, you can order requests to better cache (and evict) KV".
  • A revised cascade attention. This ordering requires a minor revision of Hydragen-style cascade attention, the technique that shares computation across requests with a common prefix.
  • Reduced host overhead. Large numbers of small requests to small models can yield low GPU utilization because the host cannot feed the GPU fast enough. When the structure of the requests is known in advance, this overhead can be avoided, the authors write.

The point, then, is not a new model or new hardware, but that the query planner's knowledge of the workload makes the inference engine's scheduling problem substantially easier.

The numbers — as their own claims

All the performance figures below come from Modal/CMU's own blog post and their own benchmark. No independent replication exists in the available coverage.

  • Over 1 billion tokens per minute per H100. On one multi-join query where scheduling matters especially much, the authors claim Quail achieves over one billion processed tokens per minute (TPM/GPU) — over ten times faster than their vLLM baseline on the same hardware. This is thus the top-line figure on a constructed query, not an average.
  • 1.84x geometric mean. On the newly launched AI-SQL benchmark, they claim that Quail is overall 1.84 times faster than vLLM, the geometric mean across the tasks. Note that the benchmark deliberately includes two queries designed to surface areas for improvement in AI-SQL inference — the authors themselves say they are made to reveal where the engine struggles. That means the average figure rests partly on unfavorable cases, making it hard to interpret without per-task results, which are not detailed in the announcement.
  • Under 6 cents per billion tokens. Run on Modal, the top performance is claimed to correspond to a cost of under 6 cents per billion tokens. This figure combines the authors' own performance measurement with Modal's own price list — it is a self-measurement on their own platform, not a market price.

It is worth noting that the striking tenfold figure applies to a single query chosen precisely because scheduling matters a lot there, while the more conservative 1.84x figure is the broader result.

What you can test yourself

The release is set up for reproduction: the quail-engine package (0.1.0) is installable via Modal, and the demo runs qwen3-4b-fp8 on one H100. The benchmark for AI-SQL queries has also been published alongside the release, so others can run the comparison against vLLM or other engines on equal terms.

Caveats and open questions

Several factors mean the figures should be treated with caution for now:

  • Self-reported benchmark. All results were measured by the developers themselves, on their own benchmark, against a baseline (vLLM) configured by them. No independent replication exists.
  • The benchmark's composition is unclear. The announcement does not give per-task results or a complete composition, and two queries are deliberately designed to expose weaknesses. That makes it hard to assess how robust the 1.84x figure is.
  • Model and workload choices. The demo uses a small model (qwen3-4b-fp8), and the mechanisms — particularly host overhead on small requests — are most relevant to exactly such workloads. How well the results generalize to larger models or other AI-SQL patterns is not documented in the announcement.
  • The cost figure is a self-measurement. "Under 6 cents per billion tokens" assumes Modal's platform and prices, and follows directly from the performance figure.

Several questions remain open: how Quail works together with established SQL engines in practice, how it scales beyond the demo, and whether independent parties can reproduce the tenfold figure. Since the benchmark and the packages are publicly available, these are questions that should be answerable through replication in the time to come — until then, it is reasonable to read the release as a promising, but undocumented-by-others, result from a team that knows both its own strengths and weaknesses.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Get the best of AI MAG in your inbox

News, analysis, and ideas at the intersection of AI and society.