Liquid AI launches d1-3B: decision model that answers calibrated in a single forward pass

On October 7, 2026, Liquid AI released two open decision models, d1-3B and d1-omni-600M, which deliver calibrated answers in a single forward pass — no generated tokens, no sampling.

Illustration: a single taut wire passing straight through a stack of glass layers, next to a pile of tangled threads — one direct path instead of many.
Illustration
Gift article

Liquid AI launches d1-3B: decision model that answers calibrated in a single forward pass

On October 7, 2026, Liquid AI released two open decision models, d1-3B and d1-omni-600M, which deliver calibrated answers in a single forward pass — no generated tokens, no sampling. The company's own measurements put responses at 8–50 milliseconds on edge hardware. But every number is self-reported, and the audio model is in practice uncharacterized.

An inverted generative pipeline

Most language models answer questions by generating one token at a time, with sampling and probably some variation from run to run. Liquid AI's d1 family turns this pipeline upside down: according to the company, the decision models do not produce tokens at all — they return an answer in a single forward pass over the network (Liquid AI).

Answers come in three fixed types, as MarkTechPost summarizes with reference to the company (MarkTechPost):

  • noul: a yes/no question, returned as P(yes) — a calibrated probability.
  • choice: a labeled choice, delivered with the full probability distribution over the options.
  • score: a probability-weighted position on an ordered rubric with 2 to 10 levels.

Multiple questions can be shared over one encoder state in a single call, making it possible to ask many decision questions about the same input in one operation. It is worth emphasizing what this is not: neither checkpoint is a chat model. They cannot answer open questions or generate text — they classify, rank, and score.

The two checkpoints

d1-3B is the main model. It has 3.12 billion parameters, a 400M SigLIP2 NaFlex vision encoder, and a 32,768-token context, and was trained from LFM2.5-VL-3B, a decoder-only vision-language model that accepts text and images as input (Liquid AI; MarkTechPost).

d1-omni-600M is the small, experimental model. With 587 million parameters, a 17-layer FastConformer audio encoder, a 16,384-token context, and audio clips of up to 30 seconds, it was trained from LFM2.5-Encoder-350M, a bidirectional encoder augmented with vision and audio encoders. It handles both text+image and text+audio.

The numbers — with caveats about who measured them

All the performance figures below come from Liquid AI itself, which on its own initiative ran the official Decision Index scorer rather than submitting results to an independent leaderboard (MarkTechPost). That does not mean the numbers are wrong, but that they currently lack independent verification.

On Decision Index v0.2.1 (public split), d1-3B scores 48.57 — according to the company ahead of every model under 10 billion parameters and on par with Decider 35B-A3B, a decision model twelve times larger (Liquid AI). The model leads the Tools (74.5) and Arts (36.3) categories but falls short on Knowledge, where it only gets 23.8 (MarkTechPost).

On seven public text benchmarks, d1-3B averages 82.9, the highest in the table and ahead of Decider 4B (81.1). d1-omni-600M reaches 78.4, ahead of Decider 2B (77.1) with a quarter of the parameters, and has the highest individual scores on Civil Comments (95.8) and PAWS-X (79.5).

The latency figures for d1-3B, measured by the company, perhaps best show why this is an edge story: 8 ms on an NVIDIA GeForce RTX 4090, 16 ms on a Jetson AGX Thor, 26 ms on a Jetson AGX Orin, and 50 ms on a Jetson Orin Nano — which the company describes as fast enough for real-time decisions on the smallest edge hardware.

For d1-omni-600M, the picture is weaker: the Decision Index score of 15.95 places it far behind its big sibling, and the company describes it as an early experimental checkpoint. MarkTechPost notes that it is "an early research release with no published latency figures" — that is, there are no published latency figures at all for this model (MarkTechPost).

License and access

Both checkpoints are on Hugging Face from launch day, with day-one support in llama.cpp. They are released under the LFM Open License v1.0, which permits free commercial use for businesses with under $10 million in annual revenue (MarkTechPost). For smaller companies building agent pipelines or content moderation, this means the technology can be adopted without licensing costs.

What are the models for?

Liquid AI recommends d1 for routing, moderation, intent classification, reranking, LLM-as-a-judge scoring, agent guardrails, and visual inspection (MarkTechPost). There is a recurring pattern: these are all tasks where a generative language model is today used as a classifier — often by generating a textual answer that is then parsed. A decision model replaces that detour with a direct, typed output with probabilities, which fits when the answer feeds into further program logic and latency must be low.

Open questions

Several caveats come with this story. First: all benchmark and latency figures are the company's own measurements, and the d1 scores have not been submitted to any leaderboard. Independent verification remains outstanding. Second, d1-omni-600M is in practice uncharacterized — weak index score, no latency figures, and Liquid AI itself acknowledges that dedicated audio benchmarks for decisions are "currently [an] open problem," and that it looks forward to the community developing such benchmarks as the category matures (Liquid AI). According to MarkTechPost, the audio encoder was trained exclusively on English-language user-to-assistant requests, which limits what can be concluded about broader audio understanding. Third, the private vision split of Decision Index v0.3 is not reported in this release, so the quality of image decisions has so far only been illuminated through the backbone's ordinary vision benchmarks. Finally: d1-3B's weakness on Knowledge (23.8) is a reminder that these are specialized tools, not general models — on questions requiring factual knowledge rather than classification, they are likely the wrong tool.

What this means in practice depends on whether the numbers hold up under independent testing. If they do, d1 points to an interesting division in AI infrastructure: generative models for what must be formulated, decision models for what must be decided.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Get the best of AI MAG in your inbox

News, analysis, and ideas at the intersection of AI and society.