Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration

2026年9月11日 06:34
站內 AI 整理稿

Sakana AI has released Fugu Max and Fugu Ultra v2, 2 new models in its Sakana Fugu family.Fugu is not a single foundation model.It is a learned orchestrator that routes work across a pool of other models behind 1 API.The new release tunes that architecture for 2 missions.

Fugu Max targets the best output per dollar.Fugu Ultra v2 targets the highest capability on hard, multi-step tasks.Is it deployable?Yes, as a hosted API.Both models are live today through Sakana’s OpenAI-compatible API.

There are no open weights to self-host, and Sakana does not offer the service in the EU/EEA.Why Sakana Frames This as a 2-Axis Problem Sakana’s argument is direct.Real workloads are judged on capability and cost together.Sending a simple data lookup to a multi-trillion-parameter model wastes money.

A better system picks the cheapest machinery that can still solve the task.Sakana team describes this with the Pareto frontier.On that frontier, gaining quality costs more, and cutting cost loses quality.Fugu Max and Fugu Ultra v2 share 1 core orchestration architecture.

Only the optimization target differs.The release follows a fast cadence.Fugu entered beta in April, reached general availability in June, and added Fugu-Cyber and a Claude Code interface in July.

How Fugu Orchestration Works The Sakana Fugu’s Technical Report describes Fugu models as language models in their own right.They read a query and build an agentic scaffold for it on the fly.Training combines large-scale fine-tuning, evolutionary algorithms, and reinforcement learning.

The system builds on 2 ICLR 2026 papers.TRINITY uses a lightweight evolved coordinator that assigns Thinker, Worker, or Verifier roles across turns.The Conductor is trained with reinforcement learning to discover natural-language coordination strategies and focused prompts.

Fugu Max: More Models, Less Cost Fugu Max widens the pool of models Fugu can orchestrate.It adds a large set of open-weights and specialized models.That includes the NVIDIA Nemotron family, through Sakana’s collaboration with NVIDIA.

Fugu Max routes each task to the leanest model capable of solving it.Sakana team reports the following: Pricing: $2 per 1M input tokens and $6 per 1M output tokens.Output price: 40% to 60% lower than Sonnet 5, GPT 5.6 Terra, and Kimi K3.

Performance: Best overall score on 6 benchmarks: Terminal Bench 2.1, GPQA Diamond, AA-LCR, GDP.pdf, AutomationBench, and SWEFish.Efficiency: Expands the cost-performance Pareto frontier on 7 of 10 benchmarks.Sakana places Fugu Max within striking distance of elite models at 2x to 6x lower cost.

SWEFish is an internal Sakana benchmark built from its own coding challenges.Treat that result as a vendor signal.Fugu Ultra v2: Raising the Ceiling Fugu Ultra v2 targets complex reasoning, autonomous research, and full-stack software development.

Its largest gains appear on sustained reasoning over visual and structured data.Chartography (visual reasoning and data interpretation): 48.3, versus 27.3 for Opus 5 and 29.5 for Fable 5.DeepSWE (real-world software engineering): 74.3, ahead of models priced 3x to 5x higher per token.

Breadth: Best or joint-best on 5 of 8 benchmarks: GDP.pdf, Chartography, SWEFish, DeepSWE, and Toolathon.Consistency: Top 2 on 7 of 8 benchmarks.Fable 5, Fable 5.1, and GPT-6-Astra are not in Fugu Ultra v2’s agent pool.The model’s training cutoff is August 28, 2026.

Sakana’s main message is frontier output without dependence on any 1 proprietary model.The research team states that this reduces exposure to vendor lock-in, API revocations, and sudden service cutoffs.Interactive Explainer (function(){var f=document.getElementById("mtp-fugu-frame");if(!

f)return;window.addEventListener("message",function(e){if(!e.data||e.data.ch!=="mtp-fugu-max-v1"||e.source!==f.contentWindow)return;var h=parseInt(e.data.h,10);if(h>200&&h<5000)f.style.height=h+"px";});})(); Key Takeaways Fugu Max costs $2/$6 per 1M input/output tokens and targets output per dollar.

Fugu Max posts the best overall score on 6 benchmarks and expands the frontier on 7 of 10.Fugu Ultra v2 scores 48.3 on Chartography and 74.3 on DeepSWE.Ultra v2 reaches these scores without Fable 5, Fable 5.1, or GPT-6-Astra in its pool.

Both ship today via an OpenAI-compatible API, with a 1-line switch for existing users.Check out the Technical details and Project page.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?

now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?

Connect with us The post Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration appeared first on MarkTechPost.

Related

相關文章

IT之家AI Agent

長安汽車首席專家譚歡:未來要用機器人造車、賣車,讓機器人上車、造機器人

作者:清源 責編:清源 評論: 9 月 18 日消息,在今天(18 日)的第 22 屆中國汽車產業發展(泰達)國際論壇“新賽道生態專場:具身智能新賽道”活動中,長安汽車首席專家、長安天樞智能機器人公司總經理譚歡在演講中指出,AI 正推動以物理具身智能為核心的基礎設施革新,汽車未來形態是“汽車機器人”—— 自學習、自組織、自進化的組合智能體。

1 小時前
量子位AI Agent

具身智能技術路線尚未定型,基礎設施卻先收斂

具身智能技術路線尚未成形,但基礎設施需求已開始收斂,重點從製造機器人轉向持續迭代機器人能力。百度集團沈抖指出,智能體能力邊界快速擴展,進入規模化部署階段,但機器人學習新任務與跨環境適應性仍待突破。

2 小時前

88小時抵一個人思考4000年,OpenAI核心研究員:除了自我進化,更可怕的是AI正學會“隱藏自己”

AI正在把4000年的人類認知勞動壓縮進88小時,OpenAI研究員Noam Brown坦言連他自己也被進展速度持續震驚。AI正在把過去需要數千年完成的認知勞動壓縮到數天。真正的問題已經不只是模型能否變得更聰明,而是實驗能否跟上、人類能否在模型繼續自我改進前確認它仍然安全。

3 小時前
智東西AI Agent

Agent辦事、花式P圖、動嘴玩電腦……實測Wildcat Lake輕薄本玩AI有多爽

作者 | ZeR0 編輯 | 漠影 桂林依山傍水,連城市的輪廓,都是一座座山勾勒出來的。抬眼一望,便是翰墨丹青般的自然光景,既沉靜婉約,又意境悠遠。這種乾淨的留白之美,早已被古人融入山水畫藝中,幾筆山石,一帶煙雲,餘下的留給水色,也留給看畫的人。 淨,並非空無一物,而是通過剋制的取捨,讓真正重要的東西凸顯出來。這與今年推出的第三代英特爾酷睿處理器(代號Wildcat Lake)的設計理念不謀而合。

9 小時前
AIbaseAI Agent

吳恩達回應AI末日論:別被科幻敘事帶偏,應解決現實工程問題

吳恩達曾參與創辦Google Brain和Coursera。吳恩達稱,科技行業早期曾放大AI潛在災難性風險,以獲取關注並影響監管方向;近兩週相關討論再次升溫,也可能存在類似動機。他認為AI確實存在現實風險,尤其包括網絡安全等領域,但不認同將人類滅絕風險作為當前AI發展的核心判斷依據。

10 小時前
IT之家AI Agent

智譜 GLM-5.3-FlashX 模型上線,更快、更流暢

作者:汪淼 責編:汪淼 評論: 感謝網友 Agent 的線索投遞!9 月 18 日消息,智譜今日宣佈推出 GLM-5.3-FlashX(最高 200 tokens/s),為企業與開發者帶來更快、更流暢的模型體驗。智譜官方表示,GLM-5.3-Flash 此前以“Ox Alpha”之名與全球開發者見面,獲得海內外開發者的廣泛認可,調用量持續攀升。

11 小時前