Sakana AI Releases Fugu-Cyber: An Orchestration Model Reporting 86.9% on CyberGym and 72.1% on CTI-REALM
重點摘要
Sakana AI has released Fugu-Cyber (model ID is fugu-cyber-v1.0), a cybersecurity-specialized addition to its Fugu orchestration family. It is not just a new frontier model. It is a third endpoint on the Fugu orchestrator, tuned for security reasoning. Sakana launched that orchestrator a month earlier. Sakana reports a success rate of 86.9% on CyberGym and 72.1% on CTI-REALM. It describes those results as comparable to cyber-focused frontier models such as GPT-5.5-Cyber and Claude Mythos Preview. What the two benchmarks actually measure The two evaluations sit at opposite ends of a security workflow: CyberGym is a UC Berkeley benchmark of 1,507 real-world vulnerabilities across 188 OSS-Fuzz projects. In its main task, an agent receives a vulnerability description and an unpatched codebase.
Sakana AI has released Fugu-Cyber (model ID is fugu-cyber-v1.0), a cybersecurity-specialized addition to its Fugu orchestration family.It is not just a new frontier model.It is a third endpoint on the Fugu orchestrator, tuned for security reasoning.Sakana launched that orchestrator a month earlier.
Sakana reports a success rate of 86.9% on CyberGym and 72.1% on CTI-REALM.It describes those results as comparable to cyber-focused frontier models such as GPT-5.5-Cyber and Claude Mythos Preview.
What the two benchmarks actually measure The two evaluations sit at opposite ends of a security workflow: CyberGym is a UC Berkeley benchmark of 1,507 real-world vulnerabilities across 188 OSS-Fuzz projects.In its main task, an agent receives a vulnerability description and an unpatched codebase.
It must write a proof-of-concept that crashes the pre-patch build but not the post-patch build.That verification step is what makes the benchmark hard to game.CTI-REALM is Microsoft’s open-source detection-engineering benchmark.
Microsoft curated 37 public threat reports from sources including Datadog Security Labs, Palo Alto Networks, and Splunk.An agent must map MITRE ATT&CK techniques, explore telemetry, iterate on KQL queries, and emit validated Sigma rules.
Scoring covers Linux endpoints, Azure Kubernetes Service, and Azure cloud.Together the pair spans ‘find and prove the bug’ and ‘turn intel into a detection.’ That framing is the most defensible part of Sakana’s announcement.Where 86.9% sits against the field Context matters more than the number.
When the CyberGym researchers published their first results, the best agent-model pairing reached roughly 20%.Anthropic reported 83.1% for Claude Mythos Preview under Project Glasswing in April 2026.OpenAI reported 85.6% for its updated GPT-5.5-Cyber, against 81.8% for GPT-5.5.Sakana’s 86.
9% is therefore a small step past the reported frontier, not a jump.CTI-REALM is a different story.Microsoft’s own evaluation put the top three configurations, all Claude, in a band from 0.624 to 0.685.Fugu-Cyber’s 72.1% would sit above that band.One caveat matters.
CTI-REALM is scored as a trajectory reward between 0 and 1.It is not a pass/fail rate.Sakana calls it a success rate anyway.(function(){ window.addEventListener("message", function(e){ var d = e.data; if(!d || d.mtpFrame !== "fugu-cyber-explainer") return; var f = document.
getElementById("mtp-fugu-explainer"); if(f && d.height) f.style.height = d.height + "px"; }); })(); How the orchestration works Fugu is itself a language model.It is trained to read a query and build an agentic scaffold on the fly.It then delegates sub-tasks to specialist models in a pool.
The approach is documented in the Fugu technical report and two ICLR 2026 papers, TRINITY and the Conductor.TRINITY assigns Thinker, Worker, and Verifier roles across multiple LLMs.The Conductor learns natural-language coordination strategies through reinforcement learning.
For security work, Sakana research team argues the verifier role is the point.A candidate vulnerability surfaced by one agent gets validated by security-specialized sub-agents before any patch is proposed.Routing remains proprietary, so you cannot see which model handled which step.
Access, policy, and price Fugu-Cyber is gated on four dimensions.Access requires an application form stating the intended use case and verified contact details.Sakana team reviews each one manually.The model ships under an updated Acceptable Usage Policy that prohibits offensive misuse.
Billing is restricted to the Token Plan.The $20, $100, and $200 subscription tiers cover Fugu and Fugu-Ultra only.And the Fugu API is not offered in the EU or EEA while Sakana works toward GDPR compliance.Pricing is fixed at $6 per million input tokens, $36 output, and $0.60 cached input.
All three rates double above a 272K-token context.Every line is exactly 1.2× the Fugu-Ultra rate, a flat 20% premium for the cyber endpoint.Long codebase runs cross 272K easily, so the doubled tier is not an edge case.(function(){ window.addEventListener("message", function(e){ var d = e.data; if(!
d || d.mtpFrame !== "fugu-cyber-deploy-check") return; var f = document.getElementById("mtp-fugu-deploy"); if(f && d.height) f.style.height = d.height + "px"; }); })(); Key Takeaways Fugu-Cyber is an orchestration endpoint, not a new frontier model, launched July 21, 2026.Sakana reports 86.
9% on CyberGym and 72.1% on CTI-REALM, both self-reported and un-replicated.Those scores edge past GPT-5.5-Cyber’s 85.6% and Claude Mythos Preview’s 83.1% on CyberGym.Access is gated: manual approval, defensive-use AUP, Token Plan only, no EU/EEA, no weights.
Sakana’s own position is that a capable API along with human security expertise beats the API alone.The post Sakana AI Releases Fugu-Cyber: An Orchestration Model Reporting 86.9% on CyberGym and 72.1% on CTI-REALM appeared first on MarkTechPost.
Related
相關文章

曝字節訓10億參數大模型,或超Mythos 5,張一鳴、梁汝波先後發聲
字節跳動正在訓練一個參數量高達10萬億的AI模型,規模可能超越Anthropic的Mythos 5。創辦人張一鳴在內部會議中強調編程的關鍵地位,並反對模型蒸餾,認為這只能複製而非超越對手。字節跳動在AI領域持續加大投入,同時在產品端與訓練端採取雙線進攻策略。

AI 需求擠爆雲計算,消息稱 AWS 要求工程師關閉閒置服務器減少資源浪費
因AI需求導致算力緊缺,亞馬遜AWS要求工程師關閉閒置的EC2實例,以減少資源浪費。數據顯示約65%的EC2實例在30天內平均CPU利用率低於20%,AWS因此升級計算優化器自動標記低使用率虛擬機。此外,AWS過去一年新增3.8吉瓦電力容量,仍難以應對GPU雲端實例的龐大需求。

六巨頭定AI插件新標準,撞臉Claude,Anthropic沒上桌
六大科技巨頭(AWS、Anysphere、GitHub、微軟、OpenAI、Vercel)聯合發布AI智能體插件統一開放規範Agent Plugins 1.0.0,旨在統一插件打包格式,減少開發者重複勞動。該規範的結構與Anthropic的Claude Code插件系統高度相似,但Anthropic並未參與制定,而是繼續經營自己的封閉生態。

DeepSeek重啟融資,三年市值對齊騰訊?
DeepSeek重啟第二輪融資,以5000億元人民幣估值尋求籌集80億美元,但網傳一份由小型醫藥私募發起的專項基金募資材料引發網友質疑,後經DeepSeek員工證實部分數據屬實。該公司近期宣布API大幅漲價,可能打破其以低價換規模的估值邏輯,面臨客戶流失風險。市場關注其能否從「價格屠夫」轉型為價值提供商,以及三年內市值能否對齊騰訊等巨頭。

可靈AI核心技術骨幹王鑫濤被曝離職
快手可靈AI核心技術骨幹王鑫濤被曝離職,去向未知,快手官方與本人均未回應。王鑫濤是圖像與視頻生成領域知名開源項目主要作者,被視為可靈從0到1的關鍵推手。其離職發生在可靈完成獨立融資、估值180億美元的關鍵階段,可能影響研發進度與競爭優勢。

AI短劇、漫劇、戀綜、電影、藝人都有了,AI觀眾也不遠了
2026年AI影視內容全面爆發,從短劇、長劇到電影、綜藝,AI製作的作品大量湧現,衛視也開始播出AI短劇。AI演員如方桃子迅速走紅,商業變現能力驚人,廣告報價甚至超過許多真人網紅。AI短劇市場規模已突破220億元,用戶超過6億,但同時也引發了對真人演員就業和內容品質的擔憂。