Pokee AI 推出 Pokee-Isaac 28B:百萬 Token 上下文長度的代理模型,專為客戶內部部署設計
重點摘要
長時程代理在累積上下文的速度上,遠快於解決任務的效率。每項工具輸出、觀察結果及中間推理步驟都會保留在上下文視窗中,而兩個關鍵能力——保留上下文並在其中保持連貫——目前幾乎僅能透過雲端端點實現。這排除了受監管行業、公共部門機構及裝置端應用,因為這些場景的資料完全不允許離開內部邊界。Pokee AI 發布了 Pokee-Isaac 28B,這是一款擁有 280 億參數、僅支援文字的基礎模型,具備 1000 萬 Token 的上下文長度,專為在客戶內部邊界內運行而設計。Pokee 研究團隊聲稱,在 RULER 基準測試中,該模型在 1000 萬 Token 長度下達到 93.3% 的準確率,與代理基準測試中最具成本效益的雲端基準模型表現持平,且可單靠一張 GPU 完成部署。
Long-horizon agents accumulate context faster than they resolve tasks.
Every tool output, observation, and intermediate reasoning step stays in the window, and the two capabilities that matter — holding that context and staying coherent across it — have so far been available almost exclusively from cloud endpoints.
That excludes regulated industries, public-sector institutions, and on-device applications, where the data is not permitted to leave the boundary at all.Pokee AI released Pokee-Isaac 28B, a 28B text-only foundation model with a 10M-token context window, designed to run inside that boundary.
The Pokee research team claims 93.3% on RULER at 10M tokens, parity with the strongest cost-optimized cloud baselines on agentic benchmarks, and a serving profile that fits a single GPU.Is it deployable Yes — but licensed, not open-weight.
Pokee AI serves Isaac through an OpenAI-compatible developer API, and licenses it for deployment inside a VPC, on-premises, or on-device.The launch announcement advertises Day-0 support for vLLM and SGLang, and single-GPU serving starting from an RTX 4090 or equivalent.
The research team publishes measurements only from a single B200-class GPU, so treat the consumer-GPU claim as vendor guidance rather than a reported result.
Company level: This fits organizations that already own their inference stack — mid-size and enterprise teams with a platform group, plus device OEMs.A solo practitioner without on-prem hardware should use the hosted API instead; the boundary argument only pays off if you have a boundary.
Industries: Healthcare and payors, financial services and insurance, defense and public sector, legal and e-discovery, and pharma or semiconductor R&D.The common trait is a rule that the data cannot cross an external API boundary, not a preference for privacy.
Applications: Whole-repository code review, multi-year contract and claims analysis, incident forensics over full log archives, and long-running tool agents that never need summarization or context pruning.
The research paper makes this second point explicitly: when enough usable context is available in-boundary, memory hierarchies and compression become optional rather than required.window.addEventListener("message",function(e){ if(e.data && e.data.pkHeight){ var f=document.
getElementById("pokee-isaac-explainer"); if(f) f.style.height=e.data.pkHeight+"px"; } }); Long-context results On RULER, Isaac stays above 93.3% at every tested length, ending at 93.3% at 10M.GPT-5.6 Luna and Gemini 3.5 Flash Lite track it to 512K, then hit context-overflow at 1M.
On MRCR v2 with 8 needles, Isaac scores 0.607, 0.743, and 0.500 at 256K, 512K, and 1M.Its margin over Gemini widens from 0.133 to 0.295 across that sweep.Agentic and security results Isaac leads BFCL v4 at 70.94 against Luna’s 70.61.
The report calls that parity rather than a lead, which is the correct read.On τ³-bench it averages 0.662 across four domains, ahead of Gemini’s 0.631, with banking at 0.186 for everyone’s difficulty.On MCP-Atlas it places third at 74.59% coverage, but uses 9.10 turns per task against Gemini’s 14.99.
On Terminal-Bench 2.1 it resolves 56 of 86 text-compatible tasks (65.1%), behind Luna’s 60.That is the one benchmark a cloud baseline wins, and the report states it plainly.On DTAP red-teaming, Isaac records the lowest direct (36.0), indirect (35.2), and combined (35.
6) attack success rates, with 82.5 benign success.One condition differs: baselines ran under the stock runner, Isaac under the Pokee harness.Efficiency, pricing, and portability Under the RULER workload on one B200-class GPU, TTFT is 23.6s at 1M and 72.9s at 10M.
Prefill throughput rises with context, from 42,400 to 137,200 tokens/s, so a ten-fold longer prompt costs roughly three times the TTFT.List pricing is $0.15/$1.00 per million input/output tokens, marked provisional.
Isaac also runs fully on-device on Intel Arc Pro B70 and Core Ultra Series 3 (Panther Lake), and on Qualcomm Snapdragon X2 Elite.Key Takeaways Pokee-Isaac 28B scores 93.3% on RULER at 10M tokens; every baseline in its panel returns 0.0 beyond 2M.
Prefill reaches 137,200 tokens/s at 10M context on one B200; decode holds flat near 335 tokens/s.It leads BFCL v4 (70.94) and τ³-bench (0.662 avg), places second on Terminal-Bench 2.1, third on MCP-Atlas.Lowest combined attack success rate on DTAP (35.6) while keeping 82.5 benign task success.
Weights are not published; deployment is licensed into VPC, on-premises, or on-device.Check out the Blog and Paper.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?Connect with us The post Pokee AI Releases Pokee-Isaac 28B: A 10M-Token Context Agentic Model Built to Run Inside the Customer Boundary appeared first on MarkTechPost.
Related
相關文章

重磅!蘋果國行AI突然發佈,首次官宣牽手阿里,Mac用上千問了
蘋果官方更新Mac使用手冊,首次公開確認與阿里巴巴千問合作,將AI能力整合進國行Apple智能。用戶可透過Siri與寫作工具直接在Mac上使用千問的文本理解與生成功能,合作僅限中國大陸地區。

【數智周報】張一鳴:字節跳動“拒絕蒸餾”,不用別人輸出換榜單排名;三星發佈下一代AI存儲路線圖,展示zHBM和400層以上V10 NAND技術;閃迪第四財季營收超預期增長372%,數據中心收入增近13倍,擬豪擲140億回購
字節跳動創辦人張一鳴在內部會議強調公司堅持長期主義,拒絕蒸餾他人模型以換取榜單排名。三星發佈下一代AI存儲路線圖,展示zHBM與400層以上V10 NAND技術。閃迪第四財季營收年增372%,數據中心收入成長近13倍,並計劃回購140億美元股票。

專為 AI 智能體打造的雲端瀏覽器,Cloudflare 發佈 Kitesurf
Cloudflare 推出專為 AI 智能體打造的雲端瀏覽器 Kitesurf,運行於 Cloudflare Workers 平台,支援瀏覽網站與填寫表單。該瀏覽器較 Chromium 消耗更少計算資源,能降低運作成本,目前處於測試階段。

都學壞了!奧特曼親手封鎖最強模型Astra,重蹈Mythos覆轍
OpenAI 執行長奧特曼宣布延後發布最強模型 Astra,原因是內部評估顯示其網路安全能力可能達到「關鍵」級別,需進一步確保安全。此舉被認為與先前 Anthropic 推遲 Mythos 模型發布的策略類似,引發外界質疑是否為恐懼行銷。

微軟預告 Win10/Win11 經典版 Outlook 更新:基於上下文 AI 解釋用戶選中文本
首頁 > 智能時代>人工智能 微軟預告 Win10/Win11 經典版 Outlook 更新:基於上下文 AI 解釋用戶選中文本 2026/8/8 11:39:29 來源:IT之家 作者:故淵 責編:故淵 評論: IT之家 8 月 8 日消息,科技媒體 Windows Latest 昨日(8 月 7 日)發佈博文,報道稱面向 Windows 10、Windows 11 經典版 Outlook 應用,微軟開始推送名為 Explain This 的 Copilot AI 功能。功能方面,用戶打開郵件後,只需選中任意文本,Copilot 就會結合整封郵件的內容,給出貼切的上下文解釋。微軟在更新說明裡寫道,這項改進讓用戶對 Copilot 有更高的控制權,可以決定其在何時、如何協助完成任務,從而在工作流中整合生產力、清晰度和決策能力。時間表方面,微軟預計經典版 Outlook 的 Explain This 功能會在 9 月底前完成推送,大多數用戶會在 9 月初看到它。 投訴水文 我要糾錯 下載IT之家APP,簽到賺金幣兌豪禮 相關文章關鍵詞:Outlook,AI,微軟微軟 Win10/Win11 新版 Outlook 收件箱塞入 Copilot,AI 摘要過去 8 小時郵件Win10/Win11 新版 Outlook 本月將集成 Planner 任務管理,微軟希望藉此推動經典版用戶遷移微軟承認 Outlook 存在嚴重 Bug,郵件中的圖片無法正常顯示微軟修復 Outlook 奇怪 Bug:打開 Office 文檔顯示空白或“已損壞”微軟增強 Outlook AI 功能:自主梳理收件箱、調整日程安排等部分微軟 Outlook 365 用戶反饋安裝 2603 更新後出現“幽靈滾動”

A股不賺API賺,梁文鋒永遠不虧
A股不賺API賺,梁文鋒永遠不虧超聚焦2026.08.08 10:57 · 來自湖北全文4149字00:00 / 12:06Alpha沒了?API有了!文 | 超聚焦最近,梁文鋒的兩門生意,走出了不一樣的曲線。私募排排網統計數據顯示,截至2026年7月31日,梁文鋒旗下幻方量化對外展示的9只主流量化產品中,8只年內收益由正轉負,僅九章幻方中證500量化多策略2號勉強維持0.04%微弱正收益。