SpaceXAI Releases Grok 4.6: A 500K-Context Frontier Model Tuned for Long-Running Agents, Coding, and Knowledge Work
SpaceXAI just released Grok 4.6.The release is a post-training upgrade over Grok 4.5 rather than a larger base model.
SpaceXAI held the foundation constant and spent the improvement on a longer supplemental training run, regenerated supervised fine-tuning trajectories, and reinforcement learning in agentic environments.Agents that stay on a task across many steps without drifting.Grok 4.
6 scores 61 on the Artificial Analysis Intelligence Index, up five points from Grok 4.5 and tied with GPT-5.6 Sol Max.The model takes 500,000 context tokens, is live today in Cursor and Grok Build, and adds a new xhigh reasoning-effort level above the ladder Grok 4.5 shipped with.Is it deployable?
Yes, in production, with a bounded set of workloads.The model is generally available through the xAI API as grok-4.6, is the default model in Grok Build, ships in Cursor on all plans, and is routable via OpenRouter, Vercel, and Cloudflare.
There is no open-weights release and no self-hosting path, so air-gapped deployments are out.Company stage: Seed-stage teams and indie developers can adopt it immediately, since Cursor and Grok Build need no harness work.
Mid-market engineering orgs are the strongest fit: API-only integration, mTLS authentication, batch and priority processing are documented.Regulated enterprises should stage a pilot first — the vendor’s brand history is a live procurement question in several buying committees.
Industries: Software and developer tooling, semiconductor and kernel engineering, hardware and CAD-adjacent design, financial research, and legal analysis.The training mix explicitly targeted several of these.
Applications: Repository-wide refactors, migration agents, research-and-synthesis pipelines over 500K-token corpora, first-pass application scaffolding from a product brief, GPU kernel optimization, and document-heavy knowledge work.What actually changed Grok 4.6 is not a larger base model.
SpaceXAI describes a longer supplemental training run than Grok 4.5 received, using curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, and an improved optimizer and training recipe.Grok 4.
5 was then used to regenerate supervised fine-tuning trajectories across reasoning-effort levels, agent harnesses, and domains spanning STEM, software engineering, and knowledge work, with problematic traces filtered by model-based checks.
Reinforcement learning followed in agentic environments covering knowledge work, general coding, web development, computer-aided design, and kernel optimization.
The behavioral insight is an important one: on longer trajectories, SpaceXAI reports more self-testing and verification, with the model checking its own work before moving on.That is a vendor observation from internal testing, not an independently measured result.
The model takes 500,000 context tokens, accepts text and image input with text-only output, has no stated text output limit, and carries a February 1, 2026 knowledge cutoff.reasoningeffort now supports low, medium, high (default), and a new xhigh level.
SpaceXAI did not publish a parameter count for Grok 4.6.Benchmarks: read the losses first On xAI’s launch table, Grok 4.6 (High) scores 61 on the Artificial Analysis Intelligence Index, up from 56 for Grok 4.5 and tied with GPT-5.6 Sol Max.
It leads the table on GDPval-AA v2 (1753 Elo, versus 1526 for Grok 4.5), AA-Briefcase (1577, versus 1313), and Harvey LAB.It trails on the coding rows that matter most to engineering teams.DeepSWE v1.1 lands at 65.9%, up 11.9 points generationally but behind GPT-5.6 Sol Max at 73%.Terminal-Bench v3.
0 reaches 26%, nearly double Grok 4.5’s 15.7% and still last of the four listed models.CursorBench v3.2 is 69.9%, FrontierCode v1.1 Extended is 61.3%, and APEX-Agents is 57.5%.Two things to note while evaluating.
First, the table’s bolded wins on GDPval-AA v2 and AA-Briefcase sit inside Artificial Analysis‘ published confidence intervals — they are statistical ties, not leads.Second, the comparison set excludes Anthropic’s Claude Opus 5, which currently tops that index.
The disclosed losses are the more reliable signal.Pricing and access Per the release notes, Grok 4.6 bills $2 / $0.50 / $6 per 1M tokens (input / cached input / output) below 200K prompt tokens, and $4 / $1 / $12 above that threshold.
The launch page also references a faster variant at double the price, with no separate model ID published.Grok Build and Cursor are offering 2× included usage for the first week.Teams should set a promptcache_key (or the x-grok-conv-id header on Chat Completions).
Without it, requests scatter across servers and cache hits become unreliable, so full input price applies.Interactive explainer window.addEventListener("message",function(e){if(e.data&&e.data.mtpGrok46Height){var f=document.getElementById("mtp-grok46-frame");if(f){f.style.height=e.data.
mtpGrok46Height+"px";}}}); Key Takeaways Grok 4.6 is a post-training upgrade on Grok 4.5, not a bigger base model.500K context, text and image input, and a new xhigh reasoning-effort level.Ties GPT-5.6 Sol Max at 61 on the AA Intelligence Index; trails on DeepSWE and Terminal-Bench.
Same $2 / $6 headline price, with rates doubling above 200K prompt tokens.Available today via API, Cursor, and Grok Build — no open weights, no self-hosting.Check out the FULL TECHNICAL DETAILS here.
Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?
Connect with us The post SpaceXAI Releases Grok 4.6: A 500K-Context Frontier Model Tuned for Long-Running Agents, Coding, and Knowledge Work appeared first on MarkTechPost.
Related
相關文章

阿里視頻大模型Wan3.0正式上線,行業評價“穩定、真實、有質感”
阿里巴巴影片生成大模型Wan3.0正式上線,單次可生成30秒影片,並首次支援doc、xls、ppt、pdf、md等文檔輸入。企業用戶普遍評價其「穩定、真實、有質感」,能穩定保持角色與場景一致性,並已進入短劇、影視、廣告等生產流程。即日起可於阿里雲百鍊、千問等平台體驗,標準版並推出限時7折優惠。

月之暗面第一代萬億參數多模態模型 Kimi K2.5 官宣月底結束服役
作者:歸瀧 責編:歸瀧 評論: 8 月 24 日消息,月之暗面 Kimi 官方微博今日宣佈,其第一代萬億參數多模態模型 —— Kimi K2.5 本月底即將結束服役。據此前報道,今年 1 月,月之暗面宣佈推出並開源了其最新的 Kimi K2.

消息稱知名 AI 研究員 Luke Metz 離開 OpenAI,加入 Meta 超級智能實驗室
作者:遠洋 責編:遠洋 評論: 感謝網友 華南吳彥祖 的線索投遞!8 月 24 日消息,據知情人士向 Axios 證實,知名 AI 研究員 Luke Metz 已加入 Meta 的超級智能實驗室(Superintelligence Labs)。

Anthropic 最強大模型 Fable 5 遇冷,企業用戶轉向更便宜 AI 產品
作者:遠洋 責編:遠洋 評論: 8 月 24 日消息,據英國《金融時報》報道,Anthropic 的美國客戶正在使用更便宜的替代品來替代其最強大的 AI 工具,這在其預計將實現有史以來規模最大的 IPO 之前,對其高支出的商業模式提出了質疑。

阿里雲視頻生成模型 Wan3.0 正式上線,支持單次生成 30 秒視頻、文檔輸入
作者:遠洋 責編:遠洋 評論: 8 月 24 日消息,阿里雲消息,今天,視頻生成模型 Wan3.0 正式上線。官方稱,Wan3.0 在生成時長、萬能創作、全能參考以及真實世界還原等維度全面升級,單次可生成 30 秒視頻,並首次支持 doc、xls、ppt、pdf、md 等文檔格式輸入,力求準確還原真實世界。

企業AI最後一公里:三路人馬在此交鋒
鄭敏芳 發表於 2026年08月24日 03:09 摘要:尋找自己的位置 2026年世界機器人大會現場,談到這一輪突然走紅的FDE(前線部署工程師),明略科技CEO吳明輝先把時間往回撥了十多年。“12年前我們就在非常認真地研究。”當華爾街見聞·問及FDE與傳統軟件部署有什麼區別時,吳明輝說,兩者都會進入客戶現場,但今天的FDE需要做得更深:一邊把Agent接進真實業務,一邊把現場形成的能力繼續沉澱回後臺。