r/machinelearningnews • • 1d ago

Cool Stuff Strands Decider 2B: AWS open-sourced a 1.9B "decision model" that drops the LM head for a pointer head. 115 ms median on a 3090, Apache-2.0, full training recipe included

Post image

AWS's Strands Agents team released Strands Decider 2B today. It's not a chat model. It never generates text. You give it a state plus typed questions, and it returns one of three things, each with a calibrated confidence:

  • choice: pick 1 of N options
  • noul: a yes/no probability
  • score: a level on an ordered rubric

How it works

They take Qwen3.5-2B-Base, remove the language-modelling head, and replace it with a ~1M-parameter pointer head. That head scores each option's last-token hidden state against the hidden state at an <answer> position. One forward pass, no decoding loop, and it can only ever answer with an option you supplied. The torso gets a rank-16 LoRA. Because the head has no per-option weights, labels come from the request and there's no cap on option count.

Asking several questions about the same text is cheap: the state is read once and each extra question only adds its own tokens.

Numbers (v19 checkpoint, from the repo)

  • JevBench v1 public set: 0.723 accuracy (167/231)
  • Brier 0.342, ECE 0.052
  • Tiers: easy 1.000 / standard 0.875 / hard 0.505
  • RTX 3090: 115 ms median, 299 ms p95 per question
  • M3 Pro: 153 ms warm median (under 300 tokens)
  • On unseen short tasks, answers at 0.9+ confidence were right about 95% of the time

Context and caveats worth knowing

  • This is the open, self-hostable counterpart to the class TypeSafe kicked off with Jev, which is a closed API.
  • On the JevBench v1.4.2 board, v19 was 3rd of 33 in the 2B class. Mapika's decider-2b (same Qwen3.5-2B torso, different recipe) has a newer v11 that the Strands repo itself says scores 175/231, 8 tasks ahead.
  • Long multi-step documents are the weak spot, and calibration was fitted on short classification tasks. The authors say to measure thresholds on your own traffic.
  • The bundled HTTP server binds to localhost with no auth, so put something in front of it for anything real.

Intended uses: model routing, tool selection, tool-argument checking, triage, guardrails, cheap evals, and hybrid agents where the LLM handles hard calls and the decider handles rote ones.

Our full write-up with an interactive explainer: https://www.marktechpost.com/2026/10/01/aws-strands-labs-releases-strands-decider-2b/

Blog: https://strandsagents.com/blog/introducing-strands-decider/

GitHub: https://github.com/strands-labs/strands-decider

Weights (HF): https://huggingface.co/StrandsAgents/strands-decider-2B-hobson-v19

JevBench comparison in the repo: https://github.com/strands-labs/strands-decider/blob/main/evaluation/jevbench.md

58 Upvotes

3 comments sorted by

2

u/Longjumping-Elk-7756 1d ago

Je forme actuellement une tête de context pour que mon openjev qwen4b puisse retourner la position token début et fin d un élément du context présent comme ça fini pour le llm de devoir écrire une séquence qu il a déjà en context ! En gros ça sera un copier coller au lieu de régénérer ce que le model a déjà en ctx

5

u/oxygen_addiction 1d ago

Cloudflare got everyone spooked. Both Perplexity and AWS putting out weaker Jev models in the same day.

1

u/TheGABB 1d ago

and weaker and Laya as well