AWS Strands Labs 推出開源決策模型 Strands Decider 2B:約 115 毫秒內完成選項決策
AWS Strands Labs releases Strands Decider 2B, an open source decision model.It does not generate text.It reads a state and typed questions, then returns a choice, a yes/no probability, or a score with a calibrated confidence.The model has 1.
9 billion parameters and runs locally on a CPU, a consumer GPU, or an Apple silicon Mac.Is it deployable?Yes, for local and self-hosted use.Weights are on Hugging Face under Apache-2.0, and pip install strands-decider gives a CLI and an HTTP server.The bundled server binds to 127.0.0.
1 with no authentication, so production needs your own auth layer.No hosted inference provider serves it yet.What a decision model does Decision models, also called System One models, became a category after TypeSafe AI launched Jev last month.An LLM can produce arbitrary output.
A decision model only picks between options or rates on a scale.Strands Decider supports 3 question types: choice: pick 1 of N options.noul: a yes/no probability between 0 and 1.score: a level on an ordered rubric.Every answer comes from the allowed options and carries a confidence.
The team states the model is worse than reasoning models on complex problems.It is unsuited for coding, chat, or summarization.Architecture: an LLM with its mouth removed The team starts from Qwen3.5-2B-Base and discards the language-modelling head.
A small pointer head of about 1 million parameters replaces it.That head compares the hidden state at the position against the hidden state at each option’s last token.One forward pass yields the result, with no decoding loop.The torso uses a rank-16 LoRA, and the head runs in fp32.
Label sets come from the request, so nothing caps the option count.The released checkpoint is v19.Asking several questions about one text is cheap.The state is read once, and each extra question adds only its own tokens.window.addEventListener("message",function(e){if(e.data&&e.data.
sdH){var f=document.getElementById("mtp-sd-frame");if(f)f.style.height=e.data.sdH+"px";}}); Benchmarks and latency The team measures accuracy and calibration on the public set of JevBench, a third-party benchmark for Jev-class models.Published v19 figures: JevBench v1 public accuracy: 0.
723 (167 of 231 tasks).Brier score 0.342, expected calibration error 0.052.Tier accuracy: easy 1.000, standard 0.875, hard 0.505.Latency on an RTX 3090: 115 ms median, 299 ms p95.Latency on an M3 Pro: 153 ms warm median under 300 tokens.On the v1.4.
2 board of September 25, v19 ranked 3rd of 33 in the 2B class.Excluding 3 models just over 2B, it ranked 1st of 30.The repo also flags a caveat.Mapika’s newer decider-2b v11 scores 175 of 231 on the Strands harness, 8 tasks ahead.Strands Decider was not on the newer v1.5.
4 composite board at the time of writing.Calibration is the practical win.On unseen short classification tasks, answers at 0.9 confidence or higher were right about 95% of the time.The team advises confirming or escalating below that threshold.How it compares FeatureStrands Decider 2B (v19)Jev 1.13.
0decider-2bDecision 2BDeveloperStrands Agents (AWS)TypeSafe AIMapikaFlyMy.AIAccessOpen weights, Apache-2.0Closed hosted APIOpen code or weightsOpen code or weightsBase modelQwen3.5-2B-Base + LoRA + pointer headUndisclosedQwen3.5-2B-Base + trained readoutMiniCPM5-2B + LoRA + pointer headSize1.
9BUndisclosed1.9B2.5B denseSelf-hostingYesNoYesYesFull training recipe and data list publishedYesNoNot verifiedNot verifiedJevBench public accuracy (v1.4.2 board)0.723Not in source table0.7100.
753Reported latency115 ms median (RTX 3090)70 to 500 ms (vendor)Not comparedNot compared Sources: Strands JevBench comparison, Benchmark Heaven JevBench, Jev product page.Latencies come from different hardware and harnesses, so they are not directly comparable.
Use cases and a guardrail example The team reports early success in model routing, tool selection, argument checking, triage, guardrails, evals, and hybrid agents.In a hybrid agent, an LLM makes the hard calls and the decider handles rote ones.
The repo example gates a weather tool call inside a Strands agent.A beforetoolcall intervention asks 2 yes/no questions.Are the arguments grounded in what the user said?Is calling now premature?If the agent guessed a city, it asks the user instead.From the CLI, routing “Help!
My payouts have been failing for 3 days!” across billing, sales and retail returns billing with confidence 0.768.Key Takeaways Strands Decider 2B returns choices, yes/no and scores, never text.It swaps Qwen3.5-2B’s LM head for a ~1M-parameter pointer head.v19 scores 0.
723 on JevBench public with ECE 0.052.Median latency is 115 ms on an RTX 3090.Weights, code, data list and recipe ship under Apache-2.0.Check out the technical details, GitHub repo and model weights.All credit goes to the researcher of this project.
Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?
Connect with us The post AWS Strands Labs Releases Strands Decider 2B: An Open Source Decision Model That Picks Options in About 115 ms appeared first on MarkTechPost.
Related
相關文章

填充 AI 資金彈藥:曝 OpenAI 融資再落袋 200 億美元,英偉達、軟銀和亞馬遜已出資約 90%
作者:故淵 責編:故淵 評論: 10 月 2 日消息,科技媒體 The Information 今天(10 月 2 日)發佈博文,報道稱英偉達與軟銀已分別向 OpenAI 支付最後一筆 100 億美元(注:現匯率約合 671.46 億元人民幣)投資,完成各自 300 億美元(現匯率約合 2,014.

Cloudflare 推出基於 Qwen 的開源多模態決策模型 Clef
作者:沁滄(實習) 責編:沁滄 評論: 10 月 2 日消息,當地時間 10 月 1 日,Cloudflare 推出基於 Qwen 的開源多模態決策模型 Clef(譜號),包含 Clef 與 Clef-flash 兩款模型。在 Jev 決策指數(Jev Decision Index)的基準評測中(Decision Index 0.

AI 倫理研究:DeepSeek 對男女一視同仁,美系模型卻“區別對待”
作者:故淵 責編:故淵 評論: 10 月 2 日消息,科技媒體 Wccftech 昨日(10 月 1 日)發佈博文,報道稱最新研究顯示相比較 Anthropic 的 Claude Sonnet 4.6 與 OpenAI 的 GPT-5.5,深度求索 DeepSeek 的 V4-Flash 模型能保持性別中立。

deepseek為什麼要在此時擁抱華為?
影子備忘錄2026.10.02 09:56 · 來自廣東全文4161字00:00 / 11:24這不是一次適配,這是一次搬家。文 | 影子備忘錄國慶假期前一天,DeepSeek放出了一條一則消息:正式開源面向華為昇騰算力平臺的基礎設施組件,涵蓋TileLang高級語言編譯工具、高性能計算庫與分佈式通信庫,所有組件與此前面向英偉達平臺的開源組件一一對應。如果你只把這當作一次普通的“國產適配”,那可能會錯過這件事真正的分量。

領導層洗牌後,Argon能否助谷歌重回前沿?
影子備忘錄2026.10.02 09:29 · 來自廣東全文5137字00:00 / 13:38谷歌AI的一場豪賭文 | 影子備忘錄2026年的AI賽道,用“瞬息萬變”來形容已經遠遠不夠。如果你每隔兩週沒有關注前沿模型的動態,打開排行榜時大概率會懷疑自己是不是錯過了整整一個季度。就在這樣近乎瘋狂的產品迭代節奏中,谷歌DeepMind於9月30日正式發佈了Gemini 4系列的首款旗艦模型Argon。距離上一代旗艦Gemini 3系列亮相,已經過去了將近十個月。

騰訊 WorkBuddy:Hy4 preview 夜間限免、Hy3 限免延期至 10 月 31 日
作者:沁滄(實習) 責編:沁滄 評論: 10 月 2 日消息,騰訊 WorkBuddy 於 9 月 30 日晚宣佈,決定將 Hy3 模型限免及 Hy4 preview 模型夜間限免均延期至 10 月 31 日。如果用戶已經體驗過 Hy4 preview —— WorkBuddy 在閒時夜間 (23:00 - 次日 8:00) 繼續提供免費體驗 Hy4 preview 的模型額度,截止 10 月 31 日,其餘時間正常消耗積分使用。