AWS launches Strands Decider 2B: an open decision model that points to answers instead of writing text
Amazon Web Services has released an open-source model that removes text generation from the agent decision loop entirely. Strands Decider 2B, developed by Amazon's Strands Labs and inspired by TypeSafe's Jev, scores pre-defined answer options in a single forward pass and returns a choice with confidence probabilities — in tens of milliseconds, according to AWS. The release came the same week Cloudflare launched its Clef models and OpenAI announced a Decisions API, turning a category barely three weeks old into an established infrastructure class.
What's new
Strands Decider 2B is a model of roughly 2 billion parameters, fine-tuned from a Qwen base from Alibaba. It answers bounded questions — choosing between pre-made options — with an answer and a probability, all in one forward pass. According to VentureBeat, which received AWS's announcement and three technical diagrams ahead of publication, Amazon's engineers removed the component that predicts the next word and replaced it with a small "pointer" component that scores the supplied answer options (VentureBeat, source 68c25a1c).
The model is published on Hugging Face as StrandsAgents/strands-decider-2B-hobson-v19 under the Apache 2.0 license, meaning commercial use, modification, and redistribution require no license fees. It ships with a pip package, a command-line interface, and an HTTP server — but no managed hosting service (MSN/syndicated, source 1bde2e90). The development environment behind the release, Strands Labs, is Amazon's organization for new tools and protocols for AI agents (TechCrunch, source b67db614).
The release date is reported slightly differently across coverage: TechCrunch dated its story to October 1, 2026, while other accounts say "Thursday" without a fixed date. Safe to say: released on or around October 1, 2026.
How the pointer architecture works
In a standard language model, the answer is generated one token at a time, which requires many forward passes and makes latency and responses effectively unpredictable. Strands Decider 2B cuts that process out entirely. According to the project's architecture documentation, reproduced in the syndicated report (source 1bde2e90), the pointer head compares the model's hidden-state representation of the input against each candidate's final-token representation via a dot product. A masked softmax turns the numbers into a probability distribution over the options that were provided. Everything happens in a single forward pass.
The training process was also minimally invasive: the pointer component was trained via rank-16 LoRA and adds just over one million parameters to a Qwen base of roughly two billion, according to AWS's information to VentureBeat (source 68c25a1c). The rest of the model is unchanged — what's gone is the text-generation mechanism itself.
The result is that the model can never "make up" an answer outside the list it is given. It can only rank the options it is served and say how confident it is in each one.
The performance numbers — and why they must be read carefully
AWS says, as reported by VentureBeat, that the model hits local decisions in tens of milliseconds on short tasks, and under 100 milliseconds on some common hardware and some inputs (source 68c25a1c). VentureBeat also reports specific figures from version 18 of the model: a median of 106 milliseconds and a 95th percentile of 296 milliseconds across 230 requests on a local Nvidia RTX 3090 card (source 68c25a1c). VentureBeat itself warns against presenting this as a general sub-100-millisecond guarantee.
Moreover, the figures vary across sources: the syndicated report gives a median of 115 ms for v19 on a consumer GPU, while AWS is separately quoted at around 150 ms median on an M3 MacBook. Different versions, different hardware, different medians. For teams evaluating the model, the reasonable reading is that decisions take on the order of 100–300 ms on consumer hardware — far faster than a round-trip call to a frontier model, but not a uniform sub-100 ms promise.
From hobby project to official release
The project started as a hobby for Marc Brooker, an Amazon distinguished engineer, who got the idea after seeing TypeSafe's Jev and tried building his own variant (source b67db614). The hobby project turned out good enough that it briefly held first place on the Jevbench ranking for models in its size class — enough that Amazon's engineers cleaned it up and released it via Strands Labs. Note: the Jevbench ranking is self-reported by the participants and not independently verified.
Brooker explains the motivation to TechCrunch: this class of model is, in his words, "a perfect decider for a workflow step" — that is, for a workflow step where the agent must choose what to do next based on where it is (TechCrunch, source b67db614). The point is reliability: with a closed domain of answers and a confidence score, the code calling the model can work with probabilities instead of free text.
A category crystallizing in three weeks
TypeSafe AI announced Jev on September 15, 2026 as the first commercial "System 1" model, with typed, probabilistic decisions and 70–500 ms latency — figures given by TypeSafe itself (source 2a760182). Less than three weeks later, the category looks like this:
- TypeSafe Jev: commercial, priced at $0.042 per million input tokens (The Register, source 158de919).
- Cloudflare Clef and Clef-flash: open-weight models built on Qwen bases, priced at $0.24 per million tokens — nearly six times the Jev price (source 158de919).
- AWS Strands Decider 2B: open source under Apache 2.0, free, runs locally, no hosting service.
- OpenAI Decisions API: announced at Dev Day; Sam Altman described it thus: "By focusing the model on that choice, we can make it extremely fast while keeping capabilities like image understanding, broad language support, and safety protections" (TechCrunch, source feba5a11).
Four players in under three weeks, four different business models — hosted API, open weight, open source, and frontier-vendor integration. The category is hardening into a recognized infrastructure class for agents' "System 1" decisions: fast, bounded choices that require no reasoning.
The counterarguments and the open questions
TypeSafe CEO Diogo Almeida does not see the new releases as real competition. "I get that people think it's a gold rush, but they might be underestimating the difficulty of making the models actually smart," he says (TechCrunch, source b67db614). The point is that removing text generation is the easy part; training the model to rank the right options and calibrate its confidence values credibly is the hard part.
Several things remain unresolved in the coverage. First, the performance and calibration results are essentially self-evaluations from AWS; there is no independent reproduction of either the Jevbench ranking or AWS's own measurements. Second, the version numbers are not fully consistent: the Hugging Face repo is named v19, while VentureBeat describes measurements from v18 — the figures from different sources therefore cannot be directly compared. Third: the model is only as good as the options it is given. It cannot explore, propose new paths, or explain itself — all it does is rank what someone else has defined.
What is nonetheless clear is that a major cloud vendor is now positioning itself in the category with a free, open model rather than a hosted service. For teams building agent workflows, that means the "System 1" decision layer can be set up locally with no license cost — and that competitive pressure now centers on quality, calibration, and verification, not price.

