What it is
Once an agent is running, many of the questions it faces don't need a paragraph of prose: should this step call a tool, which sub-agent gets the task, should this request be blocked. Strands Decider 2B does only that: given the current state and a typed question, it returns a probability for each option. It supports three question types: yes/no, pick one of several, and score on an ordered scale.
It comes from Strands Labs, an AWS experimental project, led by Amazon distinguished engineer Marc Brooker. The model starts from Qwen3.5-2B, swaps next-word prediction for a pointer structure that scores options, and adds a LoRA layer. The weights are on Hugging Face under Apache 2.0, with code, training recipe and evaluations on GitHub.
How to read the numbers
The official model card reports 72.3% accuracy on the public JevBench set (167 of 231 tasks). VentureBeat cites a median local latency of about 106 milliseconds on an RTX 3090, about 296 milliseconds at P95, and about 150 milliseconds for small tasks on an M3 MacBook. TypeSafe's Jev is available only as a hosted API, priced at $0.042 per million input tokens.
That still doesn't settle a winner: AWS's comparison charts leave out Jev itself, and there's no per-request cost estimate for self-hosting. The model card's own list of weaknesses is candid too: multi-step reasoning is weak, scoring tasks transfer poorly to new rubrics, and confidence calibration has only been validated on short classification tasks.
What it means
Its biggest value isn't the score but that it is reproducible and self-hostable: data stays on your machine, and you can retrain it for your own business. Teams already using a large model for yes/no calls inside agents can pick one or two high-frequency decision points and compare it with their current setup on their own data for accuracy and latency. Judgments that need multi-step reasoning should stay with a large model for now. TypeSafe CEO Diogo Almeida is unimpressed by the followers, saying they look more like "ML people wanting to implement a cool architecture."
via: Hugging Face model card, GitHub repository, VentureBeat report, TechCrunch report