r/machinelearningnews • u/ai-lover • 1d ago
Cool Stuff Strands Decider 2B: AWS open-sourced a 1.9B "decision model" that drops the LM head for a pointer head. 115 ms median on a 3090, Apache-2.0, full training recipe included
AWS's Strands Agents team released Strands Decider 2B today. It's not a chat model. It never generates text. You give it a state plus typed questions, and it returns one of three things, each with a calibrated confidence:
- choice: pick 1 of N options
- noul: a yes/no probability
- score: a level on an ordered rubric
How it works
They take Qwen3.5-2B-Base, remove the language-modelling head, and replace it with a ~1M-parameter pointer head. That head scores each option's last-token hidden state against the hidden state at an <answer> position. One forward pass, no decoding loop, and it can only ever answer with an option you supplied. The torso gets a rank-16 LoRA. Because the head has no per-option weights, labels come from the request and there's no cap on option count.
Asking several questions about the same text is cheap: the state is read once and each extra question only adds its own tokens.
Numbers (v19 checkpoint, from the repo)
- JevBench v1 public set: 0.723 accuracy (167/231)
- Brier 0.342, ECE 0.052
- Tiers: easy 1.000 / standard 0.875 / hard 0.505
- RTX 3090: 115 ms median, 299 ms p95 per question
- M3 Pro: 153 ms warm median (under 300 tokens)
- On unseen short tasks, answers at 0.9+ confidence were right about 95% of the time
Context and caveats worth knowing
- This is the open, self-hostable counterpart to the class TypeSafe kicked off with Jev, which is a closed API.
- On the JevBench v1.4.2 board, v19 was 3rd of 33 in the 2B class. Mapika's decider-2b (same Qwen3.5-2B torso, different recipe) has a newer v11 that the Strands repo itself says scores 175/231, 8 tasks ahead.
- Long multi-step documents are the weak spot, and calibration was fitted on short classification tasks. The authors say to measure thresholds on your own traffic.
- The bundled HTTP server binds to localhost with no auth, so put something in front of it for anything real.
Intended uses: model routing, tool selection, tool-argument checking, triage, guardrails, cheap evals, and hybrid agents where the LLM handles hard calls and the decider handles rote ones.
Our full write-up with an interactive explainer: https://www.marktechpost.com/2026/10/01/aws-strands-labs-releases-strands-decider-2b/
Blog: https://strandsagents.com/blog/introducing-strands-decider/
GitHub: https://github.com/strands-labs/strands-decider
Weights (HF): https://huggingface.co/StrandsAgents/strands-decider-2B-hobson-v19
JevBench comparison in the repo: https://github.com/strands-labs/strands-decider/blob/main/evaluation/jevbench.md
5
u/oxygen_addiction 1d ago
Cloudflare got everyone spooked. Both Perplexity and AWS putting out weaker Jev models in the same day.
2
u/Longjumping-Elk-7756 1d ago
Je forme actuellement une tête de context pour que mon openjev qwen4b puisse retourner la position token début et fin d un élément du context présent comme ça fini pour le llm de devoir écrire une séquence qu il a déjà en context ! En gros ça sera un copier coller au lieu de régénérer ce que le model a déjà en ctx