Google DeepMind 發表 Gemini 4 Argon,百萬級輸出 Token 支援程式開發、知識工作與網路防禦
Google DeepMind has just announced Gemini 4 Argon, its new frontier model and the first model of the Gemini 4 generation.It targets long-horizon software engineering, enterprise knowledge work in legal and finance, and cybersecurity defense.The biggest technical change is output length.
Argon can generate up to 1M tokens in a single response, up from 64K on earlier Gemini models.What Google Announced Google DeepMind described Argon as built for complex workflows across coding, enterprise knowledge work and cybersecurity defense.Google is taking a phased approach.
It is participating in the U.S.government’s voluntary process for pre-release model access.It will gather feedback from early testers and iterate on guardrails before a wider release.Pricing is already public.Argon launches at an introductory $2 per 1M input tokens and $10 per 1M output tokens.
Cached input tokens get a 95% discount, which works out to $0.10 per 1M.After the introductory period, pricing moves to $4 input and $20 output.Logan Kilpatrick confirmed the introductory $2 in and $10 out pricing.Why the 1M Output Limit Matters Current frontier APIs cap a single response far lower.
Claude Opus 5.5, Claude Fable 5.1 and GPT-6 Astra each allow 128K output tokens.Google team states that Argon can think deeply and generate hundreds of thousands of tokens in one trajectory.For developers, that means large refactors or long reports without splitting work across turns.
The cost is real, though.A full 1M output tokens costs $10 at introductory pricing and $20 after.Google has not disclosed Argon’s input context window.Benchmarks: Where Argon Leads and Where It Trails Google compared Argon against GPT-6 Astra, Claude Opus 5.5 and Claude Fable 5.1.
Argon leads outright on 12 of 18 benchmarks and ties for first on 1.Where it leads: DeepSWE v1.1 (long-horizon software engineering): 77.9%, a new state of the art.Opus 5.5 scores 74.2% and GPT-6 Astra 74.1%.Vals Index (economic impact across finance, coding, legal and tax): 68.9%, ranked first.
AutomationBench (Zapier, end-to-end business execution): 51.3%, ranked first.Opus 5.5 scores 42.5%.Harvey Legal Agent Benchmark: 19.6%, against 5.4% for GPT-6 Astra.LVBench (long video understanding): 91.7%, a new state of the art.Where it trails: FrontierSWE v2: 55.0%, behind GPT-6 Astra at 65.5%.
Terminal-Bench 4.0: 57.4%, behind Claude Opus 5.5 at 66.4%.OSWorld-2.0 (computer use): 69.2%, behind GPT-6 Astra at 72.6%.Artificial Analysis reported that Argon equals GPT-6 Astra on its Intelligence Index at 60% of the cost per task, using discounted prices.
Cyber Defense: Find, Validate, Patch Google trained Argon to autonomously find, validate and patch critical software vulnerabilities.Trusted defenders and internal Google teams receive it without cyber guardrails.On CWE-bench v1, which tests vulnerability remediation, Argon ties for first at 68%.
The rival models on that leaderboard run inside their own agent harnesses.Wiz is already using Argon through its Scan for Good initiative.The model found a critical vulnerability in healthcare software used by hospitals worldwide.Google says previous frontier models had missed it.
Before broad release, Google is strengthening safeguards in 4 areas: Misuse defenses for cyber and CBRN risks, including activation monitoring, under its Frontier Safety Framework.Indirect prompt injection resistance, where Argon leads Gray Swan’s IPI benchmark.
Misalignment monitoring of chain-of-thought and actions, with the ability to stop execution.Sealed, isolated sandboxes for high-risk training and evaluations.window.addEventListener("message",function(e){if(e.data&&e.data.mtpArgonHeight){var f=document.getElementById("mtp-argon-frame");if(f)f.style.
height=e.data.mtpArgonHeight+"px";}}); Argon Inside Google Thousands of Googlers already use Argon.Google shared 4 internal results: Argon agents applied memory optimizations across data centers, freeing over 300 TiB, with 500 TiB to 1 PiB projected.
Agents replaced 32K lines of SIMD code in the libgav1 Rust port.The decoder runs 2.7x faster with identical output.Agents are migrating C/C++ codebases to Rust, up to 800K+ lines in the Fuchsia Zircon kernel.Argon beat a published quantum algorithm baseline by 40% in minutes.
Comparison: Gemini 4 Argon vs Closest Competitors FeatureGemini 4 ArgonClaude Opus 5.5Claude Fable 5.
1GPT-6 AstraDeveloperGoogle DeepMindAnthropicAnthropicOpenAIAvailabilityFairwind Program onlyClaude API and cloudsClaude API and cloudsOpenAI APIMax output per response1M tokens128K128K128KContext windowNot disclosed1M1M1.
05MInput / output price (per 1M)$2 / $10 intro, then $4 / $20$4 / $20$10 / $50$10 / $50Cached input (per 1M)$0.10 (intro)$0.20$0.25$1.00Open weightsNoNoNoNoDeepSWE v1.177.9%74.2%67.4%74.1%Vals Index68.9%67.0%65.8%63.1%FrontierSWE v255.0%62.3%56.3%65.5%Terminal-Bench 4.057.4%66.4%57.9%58.
2%CWE-bench v168% (tie)67%58%68% (tie) Sources: Google, Anthropic Opus pricing, Anthropic Fable 5.1 docs, OpenAI GPT-6 Astra docs, OpenRouter.Benchmark scores are from Google’s published comparison.GPT-6 Astra prices are its short-context tier.
Key Takeaways Gemini 4 Argon raises the output limit from 64K to 1M tokens.It leads DeepSWE v1.1 (77.9%) and the Vals Index (68.9%).It trails on FrontierSWE v2, Terminal-Bench 4.0 and OSWorld-2.0.Introductory pricing of $2 / $10 is half of Claude Opus 5.5.
Access is limited to Fairwind cyber defenders; no public release date yet.Check out the technical details.All credit goes to the researcher of this project.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?
now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?
Connect with us The post Google DeepMind Unveils Gemini 4 Argon with 1M Output Tokens for Coding, Knowledge Work and Cyber Defense appeared first on MarkTechPost.
Related
相關文章

何愷明團隊新作:看貓片就能學會ARC挑戰
何愷明團隊提出NAT-ARC,一種純視覺的ARC解題方案,不依賴語言模型,而是使用ImageNet上的自然圖像進行MAE預訓練,再遷移到抽象格子推理任務。該方法在ARC-1上達到63.4%的單模型pass@2分數,集成後提升至70.2%,逼近專用LLM系統的表現,並證明了視覺預訓練能有效破解抽象推理的scaling瓶頸。

谷歌Gemini 4突然發佈!RSI加持,GPT和Opus都讓讓
。。。 假期第一天,谷歌攜Gemini 4 Argon空降多榜單第一。 拳打Opus 5.5,腳踢GPT-6 Astra。 更誇張的還在後頭,單項任務成本最低可至1.99美元,直接是Astra費用砍半。 最高百萬Token輸出上限,面向編程、金融和法律等複雜工作流,而且劃重點,網絡安全防禦能力超牛掰。 這波等等黨要贏麻了。 不過吧,咱普通用戶現在只可遠觀,暫時還吃不上。

階躍星辰“第一梯隊”,是“自嗨”嗎?
AIX財經2026.10.01 18:22 · 來自福建全文5252字00:00 / 15:27同行各自跑出了主線,階躍的“全棧”能跑通嗎?文 | AIX財經,作者|雷晶,編輯|魏佳沉寂許久的階躍星辰,正試圖擠進大模型第一梯隊。9月20日,它發佈Step 5 Preview,原生支持文本與視覺輸入,重點面向編程、軟件工程、專業知識工作和長程Agent任務。上線初期,登錄並完成首次調用後可獲得30天的Step Plan免費使用權。幾天後,Step Plan的月度套餐一度售罄。

騰訊經銷、字節駐場、Kimi借船:FDE成了大模型的新成本?
新立場Pro2026.10.01 16:16 · 來自四川全文5421字00:00 / 16:08AI公司活成了它們最討厭的樣子。文 | 新立場Pro今年 3 月 6 日上午十點,深圳騰訊大廈樓下開始排隊。人們帶著電腦,等騰訊雲工程師幫自己安裝 OpenClaw,首批八十多人在十點開始排隊,到十一點,數百個預約號已經發完。

谷歌推出了個“做題家”:Gemini 4 Argon屠榜,但幹活差點意思
字母AI2026.10.01 16:16 · 來自北京全文3510字00:00 / 09:31消失10個月的谷歌,為什麼連Pro這塊招牌都扔了?文 | 字母AIGemini 4可算來了,連名字也換了:這次的旗艦不叫Pro,叫Argon。從谷歌公佈的成績看,Gemini 4 Argon在知識工作方面表現搶眼,多項測試超過了OpenAI和Anthropic的旗艦模型。它主打一個知識面廣,從金融、法律到數學、科學,都有不錯的表現。谷歌還確認,9月15日在LMArena上亮相、表現接近GPT-6 Astra的“3.

豪擲82億美元,芯片巨頭押注的世界模型究竟是什麼
FoST未來敘事2026.10.01 16:04 · 來自北京全文3694字00:00 / 10:23當AI開始生成“世界”,影音遊內容正在被顛覆。文 | FoST未來敘事,作者 | 蔥蔥82億美元,芯片巨頭AMD以全股票形式收購世界模型公司World Labs。AMD在收購公告裡寫道:“World Labs 的模型經驗,將幫助AMD更深入地理解工作負載如何演變,進而塑造未來的技術路線圖。”換句話說,AMD看中的,不只是World Labs今天的模型能力,更是世界模型這類前沿模型可能給芯片硬件研發帶來的“前瞻性”。