GitHub TrendingGitHub Trending 5 min read

AWS Opens Strands Decider 2B for Fast Agent Choices

AWS’s Strands Labs on Thursday, 1 October 2026, released Strands Decider 2B, a small open-source decision model for fast option picking with calibrated confidence, the Strands blog and TechCrunch reported.

PC

PromptCrates Editorial

Staff Writer

0 0
AWS Opens Strands Decider 2B for Fast Agent Choices

AWS’s Strands Labs on Thursday, 1 October 2026, released Strands Decider 2B, a small open-source decision model designed to pick among options, answer yes-or-no questions, or assign scores with calibrated confidence instead of generating free-form text, according to the official Strands Agents blog and TechCrunch. Weights ship on Hugging Face under Apache-2.0, code and training materials are on GitHub, and pip install strands-decider exposes a CLI and local server. The roughly two-billion-parameter model is built for agent workflow steps that need speed and reliability more than essay-length answers—tool gating, routing, and guardrail checks that would be expensive or slow if every hop called a frontier LLM.

What a decision model is built to do

Decision models, sometimes called System One models, surged after TypeSafe AI launched Jev earlier in the month. Unlike large language models that can emit arbitrary tokens, a decider returns only from a closed set of answers and attaches a confidence score. Strands Decider supports three question types described in the Strands blog and MarkTechPost’s technical write-up: choice (pick one of N options), noul (a yes/no probability between zero and one), and score (a level on an ordered rubric). Amazon distinguished engineer Marc Brooker told TechCrunch the class of model is a natural “what is the next thing for me to do here” step in agentic workflows, offering lower latency and potentially lower cost than invoking a full LLM for every branch.

Architecturally, the team starts from a Qwen3.5-2B torso, removes the language-modelling head, and replaces it with a pointer head of about one million parameters that scores option hidden states against an end-of-sequence position. A rank-16 LoRA adapter fine-tunes the torso. The released checkpoint is v19 after many internal iterations documented in the repository. Because answers come from a single parallel pass rather than a decoding loop, the model is fast—and explicitly worse than reasoning models on complex multi-step problems. The Strands blog states it is unsuited for coding, chatbots, or document summarization, which is the point: keep those tasks on an LLM and reserve the decider for closed decisions.

Benchmarks, latency, and the open recipe

On JevBench’s public set, Strands reports v19 accuracy of 0.723 (167 of 231 tasks), a Brier score of 0.342, and expected calibration error of 0.052, with perfect accuracy on the easy tier. On the v1.4.2 board of 25 September it ranked third of thirty-three in the 2B class and first of thirty when models just over 2B are excluded. MarkTechPost relays a caveat from the repository: Mapika’s newer decider-2b v11 scores 175 of 231 on the Strands harness, eight tasks ahead, and Strands Decider was not on the newer v1.5.4 composite board at the time of writing, so the leaderboard is already moving. Strands reports a median of about 115 milliseconds on an RTX 3090, and MarkTechPost lists a warm median of about 153 milliseconds on an M3 Pro for tasks under 300 tokens—fast enough to sit in front of a tool call without feeling like another round trip to a hosted API.

Open weights matter for this category. TypeSafe’s Jev remains a closed hosted API, while Strands publishes code, data inventory, and training scripts so others can reproduce or fork the approach. TechCrunch quotes TypeSafe CEO Diogo Almeida arguing that many clones underestimate how hard it is to make the models actually smart; Strands’ public recipe is an invitation to test that claim. For teams already watching open agent harnesses such as DeepSeek Harness on GitHub Trending or containment work like NVIDIA’s OpenShell-oriented safety platform coverage, Decider 2B is another building block that can run entirely on a laptop.

Practical patterns for agent builders

Early uses cited by Strands include model routing, tool selection, evaluations, guardrails, memory and context management, and policy classification. A repo example gates a weather tool call inside a Strands agent: before the tool runs, the decider answers whether arguments are grounded in what the user said and whether the call is premature. If the agent invented a city, the intervention asks for clarification instead of fetching weather for a guess. That pattern—cheap closed questions on the critical path—is where decision models earn their keep beside frontier models that still write the final reply.

Operators should note the bundled local server binds to 127.0.0.1 without authentication, so production needs a real auth layer. Calibration guidance from the team suggests treating high-confidence answers as trustworthy enough to automate and escalating below that threshold. Hybrid agents that let an LLM handle hard reasoning while a decider handles rote branches are the near-term design Strands is betting on. For open-source builders, the release is less about beating GPT-class models and more about publishing a reproducible 2B-class decision head the community can improve—exactly the kind of GitHub-visible infrastructure news that tends to ripple through agent frameworks in the weeks after launch.

Primary reporting for this article: the 1 October 2026 Strands Agents announcement by Marc Brooker, Mike Chambers, and Fabio Nonato de Paula; TechCrunch’s same-day report by Tim Fernholz; and MarkTechPost’s technical summary of architecture, JevBench scores, and latency. Anchored figures include the 0.723 public accuracy, 115 ms RTX 3090 median latency, Apache-2.0 licensing, Qwen3.5-2B base with pointer head, and pip install strands-decider distribution.

github-trendingAWSStrandsopen-sourceagents

Related articles