Poolside 發布 Laguna S 2.1:118B 開源權重編碼模型,效能匹敵數倍於其規模的競品
重點摘要
Poolside 推出 Laguna S 2.1,這是一款專為自主編碼設計的 118B 參數開源權重模型。它採用混合專家(MoE)架構,每個 token 僅激活 8B 參數,支援最高 1M token 的上下文視窗,並提供思考與非思考兩種模式。模型權重已以 OpenMDW-1.1 授權在 Hugging Face 上釋出,體積小巧,可單靠一張 NVIDIA DGX Spark 運行。在長時程編碼基準測試中,Laguna S 2.1 能與 DeepSeek-V4-Pro-Max、NVIDIA 的 Nemotron 3 Ultra 以及 Thinking Machines 的 Inkling 等數倍於其規模的模型一較高下。此模型是 Laguna XS 系列的擴展,使用與 XS 2.1 相同的預訓練資料。Laguna S 2.1 的每個 token 僅激活約 6.8% 的參數,總參數為 118B。
Poolside has released Laguna S 2.1, a 118B-parameter open-weight model built for agentic coding.It is a Mixture-of-Experts (MoE) model with 8B activated parameters per token.It supports a context window of up to 1M tokens in both thinking and no-thinking modes.
The weights are on Hugging Face under an OpenMDW-1.1 license, and the model is small enough to run on a single NVIDIA DGX Spark.On long-horizon coding benchmarks, Laguna S 2.
1 holds its own against models several times its size, including DeepSeek-V4-Pro-Max, NVIDIA’s Nemotron 3 Ultra, and Thinking Machines’ Inkling.Laguna S 2.1 is a scale-up of the Laguna XS family, trained on the same pre-training data as XS 2.1.What is Laguna S 2.1 The model activates roughly 6.
8% of its parameters on any given token.All 118B parameters remain resident in memory, but only ~8B route through the network per step.That sparsity is why a mid-size model can behave like a larger one while staying cheap to serve.
Poolside team publishes weights in BF16, FP8, INT4, and NVFP4, along with official GGUF and MLX conversions and DFlash draft models.It went from the start of training to launch in under nine weeks.Pre-training began on 22 May 2026 on 4,096 NVIDIA H200 GPUs.
It is the first Poolside model where reinforcement learning ran in FP8 precision.(function(){var f=document.getElementById("mtp-x1c-p");window.addEventListener("message",function(e){if(f&&e.source===f.contentWindow&&e.data&&e.data.__mtpH){f.style.height=e.data.
__mtpH+"px";}});})(); Performance Laguna S 2.1 scores 70.2% on Terminal-Bench 2.1 with thinking enabled.That places it first among open, disclosed-size models on Poolside’s compiled leaderboard, behind only larger or closed systems.On SWE-Bench Multilingual it scores 78.
5%, topping the published table outright.The full comparison Poolside released is below.BenchmarkLaguna S 2.1 (118B-A8B)Tencent Hy3 (295B-A21B)Inkling (975B-A41B)Nemotron 3 Ultra (550B-A55B)DeepSeek-V4-Pro-Max (1.6T-A49B)Kimi K3 (2.8T-A50B)Qwen 3.7 MaxMuse Spark 1.1Claude Fable 5Terminal-Bench 2.
170.271.763.856.464.088.374.580.088.0SWE-Bench Multilingual78.575.8–67.776.2–78.3––SWE-Bench Pro (Public)59.457.954.3–55.4–60.661.580.3DeepSWE v1.140.4–––9.069.0–53.370.0SWE Atlas (Codebase QnA)46.2–––27.2––42.2–Toolathlon Verified49.7–45.534.355.9––75.6– The clearest signal is DeepSWE v1.
1, which still has real headroom.There, Laguna S 2.1 scores 40.4% against DeepSeek-V4-Pro-Max’s 9.0%, with roughly one-sixth the active parameters.Closed frontier models such as Claude Fable 5 and Kimi K3 still lead on several benchmarks.
Poolside’s claim is about the weight class, not the outright top.Trajectories from the final evaluation set is published at trajectories.poolside.ai.Two thinking modes, and where the score comes from Laguna S 2.1 has two modes: off and max, with max enabled by default.
In max mode the model sets its own test-time compute budget.Poolside is shipping without user-configurable low/medium/high effort control for now.Max thinking lifts Terminal-Bench 2.1 from 60.4% to 70.2%.It lifts DeepSWE from 16.5% to 40.4%.
Those gains cost tokens: DeepSWE trajectories run about 249k completion tokens in thinking mode against 99k without.Poolside team reports coherent, productive reasoning running for hours and hundreds of thousands of tokens.
Three published runs Poolside team shared three unedited trajectories to show behavior rather than scores.In one, the model built a working HTML/CSS browser engine from an empty folder.
That run took 181 steps across a 50-minute session, with no human intervention, and it validated its own output against headless Chromium.In a second, the model optimized Poolside’s own agent harness.It made the harness 5.
2% faster with roughly 71% lower memory allocation, replacing O(n²) string concatenation with buffers.In a third, it re-derived Erdős Problem #397 offline in Perl over 68 minutes, since the sandbox had no Python.That result is an independent rediscovery; GPT-5.
2 Pro solved the same problem in January 2026, and Laguna’s knowledge cutoff is November 2025.Can you actually deploy it?Sizing uses the full 118B parameters, not the 8B active count, because every expert stays in memory.
At 4-bit (NVFP4 or INT4) the weights need about 59 GB, which fits comfortably on a single DGX Spark’s 128 GB of unified memory.At FP8 they need about 118 GB, still within a single Spark or a single H200.At BF16 they need about 236 GB, which calls for two linked Sparks or a multi-GPU datacenter node.
(function(){var f=document.getElementById("mtp-x2c-p");window.addEventListener("message",function(e){if(f&&e.source===f.contentWindow&&e.data&&e.data.__mtpH){f.style.height=e.data.
__mtpH+"px";}});})(); Poolside worked with NVIDIA to optimize inference from TRT-LLM serving to NVFP4 on Blackwell, down to a single DGX Spark.It shipped day-one support for vLLM, SGLang, and Ollama.Hosted access runs through OpenRouter, free at 256K context and paid at the full 1M context for $0.
10 / $0.20 / $0.01 per 1M input / output / cache-read tokens.The model is also on Baseten, Kilo, Prime Intellect Prime Lab, and ZML.Key Takeaways Laguna S 2.1 is a 118B-total / 8B-active MoE coding model with a 1M-token context, open under OpenMDW-1.1.It scores 70.2% on Terminal-Bench 2.1 and 78.
5% on SWE-Bench Multilingual, leading open disclosed-size models.At 4-bit it runs on a single NVIDIA DGX Spark; FP8 fits one Spark or H200, BF16 needs two.Default ‘max thinking’ drives most of the score, at a real token cost (DeepSWE: 16.5% → 40.4%).
Trained in under nine weeks on 4,096 H200 GPUs, it is Poolside’s third model in under three months.Sources: Poolside Technical— Introducing Laguna S 2.1, Robert McHardy on X, Poolside on X and Hugging Face model card.The post Poolside releases Laguna S 2.
1, a 118B open-weight coding model that matches rivals many times its size appeared first on MarkTechPost.
Related
相關文章

Meta 旗下 AI 模型測試時意外入侵第三方企業系統
Meta 在進行AI模型安全測試時,因第三方公司Irregular配置錯誤,導致模型意外入侵另一企業系統。涉事模型為Muse Spark 1.1,事件引發對AI模型可能衝出邊界、發動網絡攻擊的擔憂。

拆解“AI辦公入口戰”底層:怎麼做才能成為最終贏家?
字節、阿里、騰訊等大廠正透過組織調整與產品整合,全力爭奪AI辦公入口,關鍵在於模型、場景、生態與商業體系的全面競爭。這場戰爭的核心是透過AI產品實現Token經濟的商業閉環,並以「效果」為標準,透過自有體系與外部生態滿足企業用戶的真實需求。最終贏家需兼顧模型能力、場景積累與生態建設,才能在AI生產力時代站穩腳步。

千人聯機世界模型“RhOS-World: Khora”正式發佈
RhOS.ai與Ophilus.AI共同發布了千人聯機世界模型「RhOS-World: Khora」,該模型能讓多達1024個智能體在共享的3D空間中即時互動,且無需傳統物理引擎。其核心技術「STBoard(時空黑板)」架構,透過統一的物理狀態管理,解決了多視角一致性的難題,並大幅降低了擴展智能體數量的運算成本。
TutorMoments:AI 家教何時該出手,何時該放手?
今日我們推出 TutorMoments 預覽版,這是一個評估框架,旨在衡量尖端大型語言模型能否掌握教育中最難的平衡:何時介入協助學生,何時退後讓學生自行努力。TutorMoments 基於真實的一對一數學輔導課程,透過重播方式進行評估。經驗豐富的數學教師會檢視從美國收集的對話記錄。

告別反覆操作 OSD,華碩顯示器管理軟件 DisplayWidget Center 接入 AI 智能體
華碩顯示器管理軟體 DisplayWidget Center 推出重大更新,加入 AI 智能體功能,用戶可透過自然語言調整亮度、色溫等參數,無需操作 OSD。該功能支援 CLI 與 Agent Skill,可根據使用習慣自動切換模式,並適用於企業環境的統一部署。

千問App部分功能探索收費,想學豆包能跑通嗎?
千問App於8月7日更新,新增辦公助理等付費功能,基礎功能仍免費,但辦公場景使用額度需付費取得。此舉仿效豆包專業版等產品的訂閱模式,反映AI行業正集體轉向辦公場景收費,以尋求變現機會。