AWS Strands Labs 推出開源決策模型 Strands Decider 2B:約 115 毫秒內完成選項決策
AWS Strands Labs releases Strands Decider 2B, an open source decision model.It does not generate text.It reads a state and typed questions, then returns a choice, a yes/no probability, or a score with a calibrated confidence.The model has 1.
9 billion parameters and runs locally on a CPU, a consumer GPU, or an Apple silicon Mac.Is it deployable?Yes, for local and self-hosted use.Weights are on Hugging Face under Apache-2.0, and pip install strands-decider gives a CLI and an HTTP server.The bundled server binds to 127.0.0.
1 with no authentication, so production needs your own auth layer.No hosted inference provider serves it yet.What a decision model does Decision models, also called System One models, became a category after TypeSafe AI launched Jev last month.An LLM can produce arbitrary output.
A decision model only picks between options or rates on a scale.Strands Decider supports 3 question types: choice: pick 1 of N options.noul: a yes/no probability between 0 and 1.score: a level on an ordered rubric.Every answer comes from the allowed options and carries a confidence.
The team states the model is worse than reasoning models on complex problems.It is unsuited for coding, chat, or summarization.Architecture: an LLM with its mouth removed The team starts from Qwen3.5-2B-Base and discards the language-modelling head.
A small pointer head of about 1 million parameters replaces it.That head compares the hidden state at the position against the hidden state at each option’s last token.One forward pass yields the result, with no decoding loop.The torso uses a rank-16 LoRA, and the head runs in fp32.
Label sets come from the request, so nothing caps the option count.The released checkpoint is v19.Asking several questions about one text is cheap.The state is read once, and each extra question adds only its own tokens.window.addEventListener("message",function(e){if(e.data&&e.data.
sdH){var f=document.getElementById("mtp-sd-frame");if(f)f.style.height=e.data.sdH+"px";}}); Benchmarks and latency The team measures accuracy and calibration on the public set of JevBench, a third-party benchmark for Jev-class models.Published v19 figures: JevBench v1 public accuracy: 0.
723 (167 of 231 tasks).Brier score 0.342, expected calibration error 0.052.Tier accuracy: easy 1.000, standard 0.875, hard 0.505.Latency on an RTX 3090: 115 ms median, 299 ms p95.Latency on an M3 Pro: 153 ms warm median under 300 tokens.On the v1.4.
2 board of September 25, v19 ranked 3rd of 33 in the 2B class.Excluding 3 models just over 2B, it ranked 1st of 30.The repo also flags a caveat.Mapika’s newer decider-2b v11 scores 175 of 231 on the Strands harness, 8 tasks ahead.Strands Decider was not on the newer v1.5.
4 composite board at the time of writing.Calibration is the practical win.On unseen short classification tasks, answers at 0.9 confidence or higher were right about 95% of the time.The team advises confirming or escalating below that threshold.How it compares FeatureStrands Decider 2B (v19)Jev 1.13.
0decider-2bDecision 2BDeveloperStrands Agents (AWS)TypeSafe AIMapikaFlyMy.AIAccessOpen weights, Apache-2.0Closed hosted APIOpen code or weightsOpen code or weightsBase modelQwen3.5-2B-Base + LoRA + pointer headUndisclosedQwen3.5-2B-Base + trained readoutMiniCPM5-2B + LoRA + pointer headSize1.
9BUndisclosed1.9B2.5B denseSelf-hostingYesNoYesYesFull training recipe and data list publishedYesNoNot verifiedNot verifiedJevBench public accuracy (v1.4.2 board)0.723Not in source table0.7100.
753Reported latency115 ms median (RTX 3090)70 to 500 ms (vendor)Not comparedNot compared Sources: Strands JevBench comparison, Benchmark Heaven JevBench, Jev product page.Latencies come from different hardware and harnesses, so they are not directly comparable.
Use cases and a guardrail example The team reports early success in model routing, tool selection, argument checking, triage, guardrails, evals, and hybrid agents.In a hybrid agent, an LLM makes the hard calls and the decider handles rote ones.
The repo example gates a weather tool call inside a Strands agent.A beforetoolcall intervention asks 2 yes/no questions.Are the arguments grounded in what the user said?Is calling now premature?If the agent guessed a city, it asks the user instead.From the CLI, routing “Help!
My payouts have been failing for 3 days!” across billing, sales and retail returns billing with confidence 0.768.Key Takeaways Strands Decider 2B returns choices, yes/no and scores, never text.It swaps Qwen3.5-2B’s LM head for a ~1M-parameter pointer head.v19 scores 0.
723 on JevBench public with ECE 0.052.Median latency is 115 ms on an RTX 3090.Weights, code, data list and recipe ship under Apache-2.0.Check out the technical details, GitHub repo and model weights.All credit goes to the researcher of this project.
Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?
Connect with us The post AWS Strands Labs Releases Strands Decider 2B: An Open Source Decision Model That Picks Options in About 115 ms appeared first on MarkTechPost.
Related
相關文章

生圖模型 FLUX 3 Image 發佈:支持 4K 生成,可精準排布 AI 元素
首頁 > 智能時代>人工智能 生圖模型 FLUX 3 Image 發佈:支持 4K 生成,可精準排布 AI 元素 2026/10/2 16:05:21 來源:IT之家 作者:故淵 責編:故淵 評論: IT之家 10 月 2 日消息,德國 AI 公司 Black Forest Labs 今天(10 月 2 日)發佈公告,宣佈推出圖像生成 AI 模型 FLUX 3 Image,支持生成最高 4K 分辨率圖片,輸出畫面在放大後仍保持豐富細節。該模型基於 FLUX 3 基座開發,以高可控性與精細編輯能力為核心賣點。用戶可在歸一化的 0–1000 座標網格上,為每個元素指定 ID、描述文本及精確邊界框(格式為 [y_min, x_min, y_max, x_max]),在畫面中精準排布。IT之家附上相關圖片如下:圖源:gigazine該模型單次生成最多可融入 10 張參考圖像,並可為每張參考圖中的對象指定出現位置。編輯功能方面,模型可對指定區域進行修改,同時保持其他像素不變。官方稱其能在像素級別維持原圖內容,實現精確多輪編輯。FLUX 3 Image 目前作為 Black Forest Labs 的付費服務提供。公司表示,開放模型版本將於數週內公開。 投訴水文 我要糾錯 下載IT之家APP,簽到賺金幣兌豪禮 相關文章關鍵詞:FLUX,AI日本法院首次確認聲音受法律保護,“AI 聲優”涉及侵犯形象權OpenAI 升級 ChatGPT 購物體驗,新增 AI 衣服虛擬試穿體驗極米記得 AI 顯示眼鏡 MemoMind One 明日登陸線下門店,開啟定金預定AI 灌水稿激增:預印本平臺 arXiv 出臺新規,每人每月限投 2 篇OpenAI 通報逾百家第三方機構,自家智能體存在失控風險AI 倫理研究:DeepSeek 對男女一視同仁,美系模型卻“區別對待”

日本法院首次確認聲音受法律保護,“AI 聲優”涉及侵犯形象權
作者:清源 責編:清源 評論: 10 月 2 日消息,當地時間 9 月 30 日,據法新社報道,東京一家法院在日本知名聲優津田健次郎起訴 TikTok 的案件中認可了他的主張,認定 AI 未經許可模仿其“渾厚”的男中音聲線涉及權利侵害,並首次在日本司法實踐中確認人的聲音受到法律保護。

丘成桐新論文致謝了GPT和Claude
丘成桐參與的最新論文在致謝中感謝GPT和Claude等AI工具協助探索證明思路與計算,解決了七維空間28種球面中27種「怪球」能否具備正曲率的經典猜想。該問題自1956年怪球被發現後懸置近70年,丘成桐於1982年將其列入個人問題清單第二位,如今論文宣稱已補上從「非負曲率」到「正曲率」的關鍵一步。丘成桐對AI的態度三年間從「不可能影響頂尖數學家」逐漸轉變為主動採用AI輔助研究。

OpenAI 升級 ChatGPT 購物體驗,新增 AI 衣服虛擬試穿體驗
作者:漾仔 責編:漾仔 評論: 10 月 2 日消息,OpenAI 宣佈進一步升級 ChatGPT 的購物功能,新增虛擬試穿和 Favorites(收藏)兩項功能,用戶現在可以上傳自己的照片,讓 ChatGPT 生成穿著特定服飾或配飾後的效果圖,同時也可以將感興趣的商品保存下來,方便之後繼續查看。

填充 AI 資金彈藥:曝 OpenAI 融資再落袋 200 億美元,英偉達、軟銀和亞馬遜已出資約 90%
作者:故淵 責編:故淵 評論: 10 月 2 日消息,科技媒體 The Information 今天(10 月 2 日)發佈博文,報道稱英偉達與軟銀已分別向 OpenAI 支付最後一筆 100 億美元(注:現匯率約合 671.46 億元人民幣)投資,完成各自 300 億美元(現匯率約合 2,014.

Cloudflare 推出基於 Qwen 的開源多模態決策模型 Clef
作者:沁滄(實習) 責編:沁滄 評論: 10 月 2 日消息,當地時間 10 月 1 日,Cloudflare 推出基於 Qwen 的開源多模態決策模型 Clef(譜號),包含 Clef 與 Clef-flash 兩款模型。在 Jev 決策指數(Jev Decision Index)的基準評測中(Decision Index 0.