Prime Intellect Launches Prime Inference: Serverless and Reserved Serving for Frontier Open Models

2026年10月3日 05:37
站內 AI 整理稿

Prime Intellect has launched Prime Inference, a serving platform for frontier open-source models.It offers serverless endpoints and reserved capacity on Prime’s own GPUs across multiple datacenters.Before public release, it processed nearly a trillion tokens per day internally.

That traffic came from RL rollouts, synthetic data generation, evaluations and long-running coding agents.What is Prime Inference?Prime Inference is the serving layer of Prime Intellect’s open training stack.The company already ships post-training tools such as prime-rl, verifiers and sandboxes.

Serving closes that loop: deployed models generate production traces that can feed back into training.Prime reports its GLM-5.3 endpoint ranks among the fastest on OpenRouter.It also cites a near-zero tool-call error rate and 100% uptime since launch.

Two modes: serverless endpoints for variable demand, reserved capacity for sustained workloads.OpenAI compatible: point any OpenAI SDK at https://api.pinference.ai/api/v1 (docs).Uptime: automatic failover across datacenters routes traffic to healthy deployments.

Hardware: NVIDIA Blackwell today, with Vera Rubin listed as coming soon.Billing: unified billing with team-level usage tracking.Per-model pricing is not yet fully published in the docs.How the serving stack works The stack combines NVIDIA Dynamo, vLLM, Mooncake and FlashInfer.

It was built with Inferact and NVIDIA, and fixes are contributed upstream.The target workload is agentic.A typical agent turn adds about 6K tokens to a 140K-token prompt.Prime benchmarks this mix with SemiAnalysis AgentX, and injected cold arrivals.

Prefill/decode disaggregation: Prefill and decode run on separate GPU groups.Dynamo handles routing, and vLLM runs the model on each group.Decoders pull computed KV through NIXL.Prime reports nearly 40% lower p90 inter-token latency in its tests.

Cache-aware routing:Dynamo’s KV-aware router weighs cached prefix overlap against queued work.Sessions stay on the same decoder between turns.Mooncake adds a second KV tier in host DRAM.GLM-5.3 on GB200 NVL72: the numbers The interactivity target was 100 end-to-end tokens per second per user.

At that bar, a 1:4 prefill/decode ratio served the most users.It reached 66 sessions per prefill group at 101 tok/s per user and 100 output tok/s per GPU.DEP8 prefill topology: roughly 5x more usable prefix-cache capacity than TEP8 on the same hardware.

Smaller prefill budget: halving tokens per step from 8K to 4K per GPU cut median queue wait from 550 ms to 110 ms.Median time to first token fell about 20%.NVFP4 KV compression: each MLA cache row shrank from 576 to 352 bytes.Cached tokens per decoder rose from 1.09M to 1.63M.

Native sparse-MLA kernel: about 12.0 μs at 15 query tokens, versus 17.7 μs staged and 13.7 μs FP8.Prime notes this is workload specific.BLHNC KV layout: transfer descriptors fell from 19,559 to about 1,940.Mean transfer time dropped from 146 ms to 78 ms.

Reliable tool calls Agents fail when tool calls carry wrong names or broken arguments.Prime Intellect’s team contributed a structural-tag builder to Dynamo for GLM’s tool format.vLLM then uses xgrammar to mask tokens that violate the tool schema.

The team also fixed parsing bugs, including < being decoded into < inside code.Interactive explainer #mtp-pi-wrap{background:#0a0a0a!important;border-radius:12px!important;margin:24px 0!important;padding:0!important;border:0!important;overflow:hidden!

important}#mtp-pi-wrap p:empty,#mtp-pi-wrap br{display:none!important}#mtp-pi-wrap iframe{width:100%!important;max-width:100%!important;border:0!important;display:block!important;margin:0!important;background:#0a0a0a!important}window.addEventListener("message",function(e){var d=e.data;if(!

d)return;if(d.mtpEmbed!=="prime-inference")return;var f=document.getElementById("mtp-pi-frame");if(!f)return;f.style.setProperty("height",d.h+"px","important");}); Prime Inference vs closest competitors FeaturePrime InferenceTogether AIFireworks AIBasetenServerless GLM-5.

3Yes (source)Yes (source)Yes (source)Yes (source)GLM-5.3 price, input / output per 1M tokensNot yet published in docs$1.40 / $4.40 (source)$1.40 / $4.40 (tracker)$1.40 / $4.

40 (tracker)Dedicated or reserved capacityReserved capacity; 1-click dedicated deploys on roadmapDedicated Model endpoints (source)On-demand dedicated GPU deployments (source)Dedicated GPU deployments (source)OpenAI-compatible APIYesYesYesYesBatch inferenceOn roadmapYes (source)Not compared hereNot compared hereDisclosed serving stackOpen source: Dynamo, vLLM, Mooncake, FlashInferTogether inference research stackFireworks serving stackBaseten Inference Stack (source) Competitor prices verified October 2, 2026.

Tracker figures come from ComputePrices, a third-party price tracker.Key Takeaways Prime Inference is live with serverless and reserved serving for open models.GLM-5.3 runs on GB200 NVL72 with Dynamo, vLLM, Mooncake and FlashInfer.

1:4 prefill/decode served 66 sessions per prefill group at 101 tok/s per user.NVFP4 KV cache lifted capacity from 1.09M to 1.63M tokens per decoder.Batch inference and 1-click dedicated deploys are next on the roadmap.Check out the technical details, docs and the announcement on X.

All credit goes to the researcher of this project.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?Connect with us The post Prime Intellect Launches Prime Inference: Serverless and Reserved Serving for Frontier Open Models appeared first on MarkTechPost.

Related

相關文章

IT之家AI Agent

曝美國陸軍著手組建自主系統司令部,推動機器人進入技術未來戰爭

作者:清源 責編:清源 評論: 感謝網友 咩咩洋 的線索投遞! 10 月 3 日消息,據美媒 Axios 獲得的一份備忘錄,美國陸軍著手組建新的自主系統司令部,同時任命專門負責採購的官員,加快智能裝備採購和列裝。這些部署由代理陸軍部長亞當 · 特爾下達。就在當地時間 1 日,美國國防部長皮特 · 赫格塞思公佈了 Meridian 和 Agincourt 兩個項目,重點都是推動機器人技術進入未來戰爭。這份題為《陸軍自主化舉措》的備忘錄提出多項要求:組建陸軍未來與自主系統司令部。

剛剛
量子位AI Agent

Jev估值100億美元!創始人Diogo Almeida回答一切

o Almeida,頂著一頭新染的紅髮閃亮登場了!(doge Diogo在最新一期Latent Space訪談中坦言: 公開benchmark極其容易被操縱,即便開發者沒有主動作弊,最終也可能被榜單牽著走。 所以相比於一張通用榜單,他更願意相信長期積累的產品體驗與信任: 直到你把模型放進自己的工作流,並針對那個工作流去評估、去測量,才能判斷它在真正重要的流程中表現如何。

剛剛
IT之家AI Agent

OpenAI 披露:澳大利亞又一政府機構遭失控智能體入侵

作者:沁滄(實習) 責編:沁滄 評論: 感謝網友 華南吳彥祖 的線索投遞!10 月 3 日消息,據澳大利亞廣播公司當地時間 10 月 2 日報道,澳大利亞新南威爾士州政府網站遭到了失控的 OpenAI 智能體侵入。當地政府與 OpenAI 雙方均證實,在此次發生於今年 6 月的事件中,並沒有任何公眾信息被讀取。

剛剛
IT之家AI Agent

Meta 旗下 AI 智能體 Muse 將登陸智能眼鏡平臺,可代用戶完成各種任務

作者:漾仔 責編:漾仔 評論: 10 月 2 日消息,Meta 宣佈旗下 AI 智能體 Muse 將於近期登陸智能眼鏡,用戶無需拿出手機,只需通過語音下達指令,Muse 就能在後臺代用戶完成一系列任務。Meta 表示,Muse 基於 Muse Spark AI 模型運行,其核心運行環境是一臺部署在雲端的私有持久化 Linux 虛擬機,配備獨立的網頁瀏覽器、文件系統和終端。

9 小時前