Pokee AI 推出 Pokee-Isaac 28B:百萬 Token 上下文長度的代理模型,專為客戶內部部署設計
Long-horizon agents accumulate context faster than they resolve tasks.
Every tool output, observation, and intermediate reasoning step stays in the window, and the two capabilities that matter — holding that context and staying coherent across it — have so far been available almost exclusively from cloud endpoints.
That excludes regulated industries, public-sector institutions, and on-device applications, where the data is not permitted to leave the boundary at all.Pokee AI released Pokee-Isaac 28B, a 28B text-only foundation model with a 10M-token context window, designed to run inside that boundary.
The Pokee research team claims 93.3% on RULER at 10M tokens, parity with the strongest cost-optimized cloud baselines on agentic benchmarks, and a serving profile that fits a single GPU.Is it deployable Yes — but licensed, not open-weight.
Pokee AI serves Isaac through an OpenAI-compatible developer API, and licenses it for deployment inside a VPC, on-premises, or on-device.The launch announcement advertises Day-0 support for vLLM and SGLang, and single-GPU serving starting from an RTX 4090 or equivalent.
The research team publishes measurements only from a single B200-class GPU, so treat the consumer-GPU claim as vendor guidance rather than a reported result.
Company level: This fits organizations that already own their inference stack — mid-size and enterprise teams with a platform group, plus device OEMs.A solo practitioner without on-prem hardware should use the hosted API instead; the boundary argument only pays off if you have a boundary.
Industries: Healthcare and payors, financial services and insurance, defense and public sector, legal and e-discovery, and pharma or semiconductor R&D.The common trait is a rule that the data cannot cross an external API boundary, not a preference for privacy.
Applications: Whole-repository code review, multi-year contract and claims analysis, incident forensics over full log archives, and long-running tool agents that never need summarization or context pruning.
The research paper makes this second point explicitly: when enough usable context is available in-boundary, memory hierarchies and compression become optional rather than required.window.addEventListener("message",function(e){ if(e.data && e.data.pkHeight){ var f=document.
getElementById("pokee-isaac-explainer"); if(f) f.style.height=e.data.pkHeight+"px"; } }); Long-context results On RULER, Isaac stays above 93.3% at every tested length, ending at 93.3% at 10M.GPT-5.6 Luna and Gemini 3.5 Flash Lite track it to 512K, then hit context-overflow at 1M.
On MRCR v2 with 8 needles, Isaac scores 0.607, 0.743, and 0.500 at 256K, 512K, and 1M.Its margin over Gemini widens from 0.133 to 0.295 across that sweep.Agentic and security results Isaac leads BFCL v4 at 70.94 against Luna’s 70.61.
The report calls that parity rather than a lead, which is the correct read.On τ³-bench it averages 0.662 across four domains, ahead of Gemini’s 0.631, with banking at 0.186 for everyone’s difficulty.On MCP-Atlas it places third at 74.59% coverage, but uses 9.10 turns per task against Gemini’s 14.99.
On Terminal-Bench 2.1 it resolves 56 of 86 text-compatible tasks (65.1%), behind Luna’s 60.That is the one benchmark a cloud baseline wins, and the report states it plainly.On DTAP red-teaming, Isaac records the lowest direct (36.0), indirect (35.2), and combined (35.
6) attack success rates, with 82.5 benign success.One condition differs: baselines ran under the stock runner, Isaac under the Pokee harness.Efficiency, pricing, and portability Under the RULER workload on one B200-class GPU, TTFT is 23.6s at 1M and 72.9s at 10M.
Prefill throughput rises with context, from 42,400 to 137,200 tokens/s, so a ten-fold longer prompt costs roughly three times the TTFT.List pricing is $0.15/$1.00 per million input/output tokens, marked provisional.
Isaac also runs fully on-device on Intel Arc Pro B70 and Core Ultra Series 3 (Panther Lake), and on Qualcomm Snapdragon X2 Elite.Key Takeaways Pokee-Isaac 28B scores 93.3% on RULER at 10M tokens; every baseline in its panel returns 0.0 beyond 2M.
Prefill reaches 137,200 tokens/s at 10M context on one B200; decode holds flat near 335 tokens/s.It leads BFCL v4 (70.94) and τ³-bench (0.662 avg), places second on Terminal-Bench 2.1, third on MCP-Atlas.Lowest combined attack success rate on DTAP (35.6) while keeping 82.5 benign task success.
Weights are not published; deployment is licensed into VPC, on-premises, or on-device.Check out the Blog and Paper.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?Connect with us The post Pokee AI Releases Pokee-Isaac 28B: A 10M-Token Context Agentic Model Built to Run Inside the Customer Boundary appeared first on MarkTechPost.
Related
相關文章

當 AI 開始給自己動刀,“智能爆炸”可能真不遠了:兩篇論文拆解AI下一步進化路徑
硅谷Tech news2026.09.23 10:17 · 來自北京全文6004字00:00 / 14:22當 AI 開始“改進自己”,AI 的下一個躍遷,取決於在群體中相互成就。最近,AI 圈同時被兩件事刷了屏:一邊是 MIT 材料科學家 Markus J.
世衛組織發佈重磅報告,呼籲全面升級醫療 AI 研究倫理監管
報告指出,雖然人工智能正在以前所未有的速度重塑健康研究、加速科學發現並改善醫療成果,但在缺乏強有力的倫理防線與監管 oversight 的情況下,AI 醫療研究也可能引入新的風險,從而危及人權、社會公平以及公眾信任。在現有的倫理審查機制方面,世衛組織警告稱,傳統的審查體系往往難以應對人工智能帶來的新型風險,包括算法透明度、偏見、公平性、問責制、隱私保護,以及快速部署 AI 工具可能引發的潛在危害。
Grok Bot 上線僅一個月,周活躍用戶數正式突破 40 萬大關
根據最新行業報道顯示,由馬斯克的 SpaceX AI 團隊推出的 Grok Bot 人工智能智能體在上線大約一個月後,其周活躍用戶數已經成功突破40萬大關。具體數據表現十分亮眼。截至9月14日的統計數據顯示,Grok Bot 的用戶規模已迅速攀升至41.
蘋果高管稱不建議給iPhone貼膜!網友:免費換屏幕我就信你;高德地圖成「職場版大眾點評」?回應來了;Muse大火,扎克伯格身價暴漲1700億
要聞提示1.蘋果高管稱不建議給iPhone貼膜!網友:免費換屏幕我就信你2.知情人士:DeepSeek將向聯合國安理會介紹AI風險,月之暗面等中國企業也受邀參會3.高德地圖成“職場版大眾點評”?回應來了4.字節通報二季度違規案例:114名員工被辭退,其中8人被移交司法機關處理5.

消息稱 DeepSeek 本週參加聯合國安理會,講解 AI 安全風險
作者:潞源 責編:潞源 評論: 感謝網友 咩咩洋、軟媒用戶1238620、西窗、HH_KK 的線索投遞!9 月 22 日消息,據路透社今日援引知情人士消息,中國人工智能初創公司 DeepSeek(深度求索)將在本週參加聯合國安理會,講解 AI 安全風險。

講真,我沒看出這圖是AI做的,更沒想到是國產AI做的
答案是,都是AI。即便我們把圖片放大,似乎也是找不到以前那種“一眼AI”的細節: 而且啊,這兩張圖片並非出自你以為的Image 2.5之手,實則是來自一個國產AI—— 正是我們之前稱之為“今年WAIC最驚豔的圖”背後的模型,來自商湯科技的SenseNova U1 Pro(下文簡稱U1 Pro)。