Mistral AI Releases Mistral Large 4 (Le Chonk): A 1.05T Parameter Multimodal MoE Model
Mistral AI has just announced the release of Mistral Large 4 (ML4), internally nicknamed Le Chonk, as a public preview.ML4 is a granular Mixture of Experts model with 1.05 trillion total parameters, 49 billion active per token, a 1.
6 billion parameter vision encoder, and a 1 million token context window, per the model documentation.It was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own European datacenters.
TL;DR Mistral AI released Mistral Large 4 (‘Le Chonk’) as a public preview on 6 October 2026: a 1.05T parameter granular MoE with 49B active per token, native image input, and a 1M context window, trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own EU datacenters.
The API is live now at $1.36 per 1M input and $4.18 per 1M output tokens, but the weights do not ship until end of October, so self hosting is not yet possible.
Its standout results are in cybersecurity, where Mistral reports 93% on Cybench and 82% on CyberGym-E2E and notes that several closed frontier models score near zero because they refuse the task.What is the architecture?ML4 is a hybrid instruct-and-reasoning MoE that takes image input natively.
Only about 4.7% of the weights activate per token, which is how a 1 trillion class model serves at mid-tier pricing.The full 1.05T still has to sit in memory, so the activation count sets compute, not your hardware bill.
Mistral has not yet published the expert count, top-k routing, or layer layout; those arrive with the weights.Training data spanned more than 160 languages, including every official EU language.Interactive Explainer (function(){ window.addEventListener('message', function(e){ if(!e.data || typeof e.
data.mtpML4Height !== 'number') return; var f = document.getElementById('mtpML4Frame'); if(f && e.data.mtpML4Height > 200){ f.style.height = e.data.mtpML4Height + 'px'; } }); })(); How does it perform?
In cybersecurity, Mistral reports 93% on Cybench and 82% on CyberGym-E2E, placing ML4 in the global top 5 on the Artificial Analysis Cyber Index.The more interesting claim is structural: Mistral states several frontier closed models score near zero on CyberGym-E2E because they refuse outright.
Reproducing a vulnerability to prove it is real is standard defensive work, and provider-level refusals block it.In agentic coding, Mistral reports 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, and 28.3% on Terminal-Bench 4.0, for a combined Artificial Analysis Coding Agent Index of 49.8%.
Mistral notes these were evaluated privately ahead of the harness going public, so they are not yet independently reproducible.A blind human evaluation run with Surge AI is the more honest signal.Professional annotators rated ML4 Preview 3.74 out of 5, second of 5 models, ahead of GLM-5.3 (3.
60) and Kimi K3 (3.59), but behind Claude Opus 5 at 4.22.On safety, ML4 resists 93.3% of attacks on Lakera’s B3 benchmark and scores 1.691 of a maximum 2.0 on KORABench.How does it compare to its closest open-weight rivals?FeatureMistral Large 4DeepSeek V4 ProKimi K3GLM-5.3Total parameters1.05T1.
6T2.8TNot officially publishedActive per token49B49B~104BNot officially publishedContext window1M1M1M1MNative image inputYesNoYesNoWeights availableNot yet, due end Oct 2026Yes, on Hugging FaceYes, since 27 Jul 2026Yes, per Artificial AnalysisLicenseNot yet announcedMITModified MITGLM-5.
3 LicenseAPI price per 1M in/out$1.36 / $4.18Varies by provider$3.00 / $15.00$1.40 / $4.40Released6 Oct 2026Aug 2026 (0813 build)16 Jul 202614 Aug 2026 What can you build with it today?
The preview API supports function calling, structured outputs, document QnA, batching, and the Agents and Conversations endpoints.Cached input is priced at $0.14 per 1M tokens, which materially changes the economics of long-context agent loops at a 1M window.Key Takeaways 1.
05T total parameters, 49B active per token, 1M context, 1.6B vision encoder.Trained on 3,800 Grace Blackwell GPUs in Mistral’s own EU datacenters.API preview live now at $1.36 per 1M input and $4.18 per 1M output tokens.Weights promised by end of October 2026, so self hosting is not yet possible.
Strongest results are in cybersecurity, where closed models often refuse the task.Check out the Mistral Large 4 announcement, Mistral Large 4 model docs and Artificial Analysis model comparisons.All credit goes to the researcher of this project.
Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.[Sponsored] The web is the one API most agents are missing.Databases, calendars and repos have APIs.
The open web mostly doesn’t.The TinyFish MCP server gives any MCP client four tools: TinySearch, TinyFetch (full pages as markdown, JavaScript included), TinyBrowser for logins and forms, and TinyAgent for multi-step jobs.Search and Fetch are free.
The post Mistral AI Releases Mistral Large 4 (Le Chonk): A 1.05T Parameter Multimodal MoE Model appeared first on MarkTechPost.
Related
相關文章
800億!曝DeepSeek新融資,即將IPO
編譯 | 畢偉豪 編輯|心緣 10月6日消息,據外媒彭博社今日報道,DeepSeek即將敲定至少800億元人民幣的新一輪融資,騰訊和寧德時代承諾的出資金額均位居該輪融資最高的一批。據彭博社報道,知情人士稱,根據已簽署的條款,該輪融資的最終募資總額可能接近1000億元,為計劃在2027年初進行的IPO鋪路。
超越特定領域的世界模型:JEPA-Anything 以單一方法涵蓋七大領域
來自 PhAI Labs、香港中文大學、復旦大學、史丹佛大學、牛津大學與普林斯頓大學的研究團隊釋出 JEPA-Anything,這是一個與領域無關的世界模型建構框架。它不為每個領域設計專屬預測模型,而是用一套共享學習方案應對截然不同的系統。該框架採用正交預測分解(OPF)技術擴充了聯合嵌入預測架構(JEPA)。研究團隊在七大領域進行測試:視覺、生物學、臨床軌跡、控制、分子動力學、物理場域與天氣。JEPA-Anything 解決了什麼問題?標準 JEPA(如 I-JEPA 或 V-JEPA 2)使用上下文編碼器、EMA 目標編碼器與一個預測器,而預測器只輸出單一的整體目標嵌入,研究團隊稱此為容量分配問題。
Reflection AI Introduces Beam: A 501B Open-Weight MoE Model With 23B Active Parameters for Coding and Agentic Workloads
Reflection AI has introduced Beam, its first open-weight model. Beam is a sparse Mixture-of-Experts (MoE) model with 501B total parameters and 23B active per token, built for coding, reasoning and agentic workloads.

最火AI崗位FDE:月薪5萬,都幹這些…
FDE(前線部署工程師)是近期最受關注的AI職位之一,海外年薪中位數約20萬美元,國內大廠也開出月薪三到五萬元。這份工作強調駐場梳理客戶的業務本體(Ontology),溝通時間佔七成以上,開發僅約三成。從業者認為,FDE與傳統外包不同,關鍵在於能否將經驗沉澱回自家產品並複用。
DeepSeek Harness v0.2 為其開源代理框架帶來官方桌面應用程式
DeepSeek 已為 DeepSeek Harness (dsh) 推出官方桌面應用程式,dsh 是其開源代理框架。該應用程式隨 v0.2 預覽版一同發布,安裝檔支援 macOS(Apple 晶片)與 Windows(64 位元)。目前可作為預覽版部署使用,使用者可從 deepseek.com/harness 下載,或執行 npx @deepseek-ai/dsh web。DeepSeek 提醒未來可能會有破壞相容性的變更。 v0.2 新增內容:框架是將模型轉化為代理的執行環境,能讀取檔案、執行指令並維持計畫。v0.2 預覽版針對日常工作與程式開發進行優化。內建功能包括:預載常用辦公室與開發工具;新增插件管理頁面,可安裝、設定、啟用與停用插件;以及右側邊欄提供檔案與差異審查預覽。

DeepSeek擴招!彈性計算團隊大量HC,尤其需要資深工程師
彈性計算團隊大量HC招人!尤其需要資深工程師。三週前不是剛招過一輪嗎?咋又缺人了。這次沒發崗位JD,直接甩了一篇DeepSeek技術分享: 《DeepSeek彈性計算(DSec):面向大規模Agent訓練的沙盒基礎設施》 現在DeepSeek的一套DSec擴展分片,大約有160臺服務器、3萬個CPU核心和250TB內存,每天要服務約300萬個沙盒。