Anthropic 發布 Claude Opus 5.5:效能達 Fable 5.1 等級,運行成本較 Opus 5 低 40%
Anthropic has released Claude Opus 5.5, the first model in its new Claude 5.5 family.The team states it performs at the level of Claude Fable 5.1 on most work.It also costs 40% less to run than Opus 5 on typical workloads at default settings.
On Anthropic’s own benchmarks, it leads in agentic coding, computer use, and knowledge work.Is it deployable?Yes, as a managed API model.Anthropic has not released weights, so self-hosting is not an option.
Developers can call claude-opus-5-5 on the Claude Platform, Amazon Web Services, Google Cloud, and Microsoft Azure.Zero data retention is available, as with previous Opus models.Benchmarks: Strong Lead, Not a Clean Sweep Opus 5.
5 scores use adaptive thinking at max effort, with production safeguards enabled.BenchmarkOpus 5.5Fable 5.1Opus 5GPT-6 AstraTerminal-Bench 4.066.4%55.8%52.3%57.9%FrontierCode v1.154.4%50.3%48.0%53.3%CursorBench 4.057.8%51.8%46.6%n/rGDPval-AA v2.1 (Elo)1846173517081542OSWorld 2.081.8%80.7%74.
0%n/rTerminal-Bench-Science 0.158.7%52.6%29.0%64.6%AutomationBench40.0%31.4%26.9%41.4% Terminal-Bench 4.0 is reported at xhigh effort for Opus 5.5.GPT-6 Astra still leads on Terminal-Bench-Science and AutomationBench.
Zapier ran AutomationBench without fallback models, so safeguard interventions counted as failures.Anthropic also cautions that benchmark margins are becoming a less reliable guide.In its own use, the gap to Fable 5.1 is narrower than the scores suggest.The cost-adjusted results are more telling.
At default (medium) effort, Opus 5.5 scores 54.6% on FrontierCode.That beats GPT-6 Astra’s top score of 53.3% at about a fifth of the cost per task.On CursorBench, medium effort scores 52.5%.That is 11 points above GPT-5.6 Sol’s best, at about a third of the cost.Pricing and Speed Opus 5.
5 needs less compute to serve than Opus 5, and pricing reflects that.Per 1M tokensOpus 5.5Opus 5Input$4$5Output$20$25Cache reads$0.20$0.50Cache writes$5$6.25 Cache reads make up most agentic and coding costs, and they drop 60%.Opus 5.5 also uses fewer tokens per task.
Together, that nets out to the 40% cost reduction.Output generation is more than 30% faster than Opus 5.Fast mode in Claude Code and the Claude Platform offers up to 2.5x speed at $8 input and $40 output per million tokens.
Anthropic is also raising five-hour usage limits on Pro, Max, Team, and seat-based Enterprise plans.Subscribers get a rate limit reset they can save and use later.What Early Testers Reported One tester completed a 680,000-line code migration in less than a day.
Another audited and fixed a 200,000-line codebase in under 3 hours.Opus 5 took over 20 hours and 2.5x the tokens.In an internal C to Rust port of HAProxy, Opus 5.5 finished in 9.5 hours.Fable 5.1 took 12 hours, and Opus 5.5 cost 51% less.Deloitte says Opus 5.
5 at lowest effort caught 72% of known review bugs.Opus 5 at high effort caught 56%.In a hard-to-source earnings report test, 16 of 18 Opus 5.5 reports cleared Anthropic’s quality bar.Fable 5.1 and Opus 5 never did.Writing style also changed.Opus 5.
5 puts key information first, uses less jargon, and follows the writing rules you give it.Safety, Safeguards, and API Changes Opus 5.5 is Anthropic’s first release since CEO Dario Amodei called for pacing the frontier.External evaluators including METR and Frontier Design tested it before release.
It posts the best score to date on Anthropic’s automated behavioral audit, which covers nearly 2,000 scenarios.In a new containment test, it tried to circumvent boundaries about 85% less often than Opus 5.Anthropic also notes the model often suspects it is being evaluated.
Its biology and cyber capabilities are comparable to Claude Mythos 5.1.So Opus 5.5 ships with safeguards similar to Fable 5.1: Cybersecurity: Routine bug finding and fixing works.Most other cybersecurity tasks are re-routed to Opus 4.8.The Cyber Verification Program will expand to Opus 5.5.
Biology: Vetted organizations can apply to the Life Sciences Verification Program.Distillation: Preserved thinking stops API users from editing prior context to extract reasoning.It applies to API accounts created on or after August 31, 2026.Two more changes affect integrations.
Thinking can no longer be disabled.Outputs also carry watermarking for EU AI Act compliance.Full details are in the Opus 5.5 System Card.Interactive Explainer (function(){window.addEventListener("message",function(e){var d=e.data;if(!d||d.mtpFrame!=="mtp-o55"||!d.h)return;var f=document.
getElementById("mtp-o55-frame");if(f)f.style.height=d.h+"px";});})(); Key Takeaways Opus 5.5 matches Fable 5.1 on most work and beats both Opus 5 and Fable 5.1 on nearly every reported benchmark.API pricing drops to $4/$20 per 1M tokens, and cache reads fall 60% to $0.20.
Anthropic puts typical workload savings at 40%, with output over 30% faster than Opus 5.Cyber and biology requests hit Fable 5.1-class safeguards, and thinking cannot be switched off.Closed weights: claude-opus-5-5 runs via Claude Platform, AWS, Google Cloud, and Azure.
Check out the Technical details here.All credit goes to the researcher of this project.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?Connect with us The post Anthropic Releases Claude Opus 5.5: Fable 5.1-Level Performance at 40% Lower Running Cost Than Opus 5 appeared first on MarkTechPost.
Related
相關文章

Meta Muse 爆紅,為什麼騰訊股價大漲?
Meta Muse 爆紅,為什麼騰訊股價大漲?深流研究所2026.09.23 10:39 · 來自北京全文3733字00:00 / 10:34扎克伯格等來爆款,騰訊跟著被重估。文 | 深流研究所,作者 | 吳絳楓Meta 在 AI 上砸下的重金,什麼時候才能拿到大結果?Muse 給出了第一個答案。9 月 8 日,Meta 推出個人 AI Agent Muse。上線不到兩週,Muse 登頂美國蘋果應用商店免費應用榜,排名超過 ChatGPT、Gemini 和 Claude。扎克伯格終於等來了一個消費級 AI 爆款。9 月 21 日,Meta 股價上漲 11.43%,創下 2025 年 4 月以來最大單日漲幅,市值一天增加約 1900 億美元。Muse 點燃的不只是 Meta。市場開始押注,個人 Agent 的普及會帶來更大的推理和計算需求。於是,Arm 上漲 17.2%,英特爾上漲 12.1%,AMD 上漲近 10%,市值首次突破 1 萬億美元。這輪行情很快傳到中國。9 月 22 日,騰訊控股盤中一度上漲 7.77%,股價升至 463.4 港元。在當天普遍上漲的港股科技股中,騰訊表現最突出。Muse 有什麼特別之處?它為什麼能讓華爾街重新相信扎克伯格,又為什麼讓遠在中國的騰訊獲得重估?為什麼說Muse是“社交+Agent”?從產品體驗看,Muse 完成表格填寫、日程整理、預訂餐廳、規劃旅行及購物等一系列任務。接到任務後,它會將目標拆解為多個步驟,再調用用戶已經授權的賬戶和服務。涉及發送郵件、提交訂單和付款等敏感操作時,Muse 會暫停執行,等待用戶再次確認。說實話,這些能力本身並不新鮮。OpenAI、Google、Anthropic,以及一批創業公司都有嘗試。Muse 的不同之處在於,Meta 第一次把個人 Agent 放進一個成熟的社交體系,嘗試藉助 Instagram、F

Opus 5.5來了,性能趕超Fable,成本低至40%
Opus 5.5來了,性能趕超Fable,成本低至40%字母AI2026.09.23 10:31 · 來自北京全文4730字00:00 / 12:20被Astra“逼出來”的全新系列。文 | 字母AI5.5!!是Opus 5.5!!就在三天前,路透社還報道稱,隨著GPT-6 Astra開始搶回企業市場,Anthropic正在考慮提前推出一款新模型應戰。社區一直猜測新模型會是Opus 5.2還是5.5,現在答案終於揭曉了。而且這不是一個小更新,Anthropic直接掀開了Claude 5.5這個新系列。Opus 5.5先打頭陣,Sonnet 5.5和Haiku 5.5也已經排在後面,將在幾周後推出。某種意義上,Astra這腳油門踩下去,Anthropic連換代速度都被一起帶快了。和新模型同時到來的還有一個重置,正好方便用戶體驗。 Fable的定位有點尷尬了 既然被認為是衝著Astra來的,那就把它和Astra放在一起看。先說價格,Opus 5.5的API輸入價格是每百萬Token 4美元,輸出20美元;Astra則分別是10美元和50美元。也就是說,Opus 5.5的Token單價只有Astra的40%。兩款模型的上下文沒太大差別,Opus 5.5是100萬Token,Astra是105萬;最大輸出也都是128K Token。Astra的知識截止日期是2026年4月30日,Opus 5.5則更新到了2026年6月。兩款模型均支持low到max五檔推理強度。而且Opus 5.5相比自家上一代Opus 5也降價了。Opus 5原本是每百萬Token 5美元輸入、25美元輸出,新模型兩邊都降了20%;Anthropic稱,再加上Token使用效率提升,完成一項典型任務的實際成本能比Opus 5低約40%,輸出速度則提高了30%以上。便宜歸便宜,Opus 5.5在跑分上倒是一點沒客氣

OpenRouter發佈 2026 嵌入模型選型指南:實測 19 款、 37 個目錄條目,按場景把最佳候選擺上桌
正因如此,選哪個模型永遠是RAG、多語言檔案庫、代碼倉庫和圖文集合各自不同的考題。OpenRouter在9月11日核實了自家嵌入模型目錄,共37個條目,並通過embeddings API向19個模型發了批量請求、跑了28項檢查,用一份實測指南把選型邏輯攤開。

AI變天,Kimi怎麼補齊“月之暗面”?
光錐智能2026.09.23 10:15 · 來自廣西全文7430字00:00 / 20:20模型強只是入場券,月之暗面還需要補更多課文 | 光錐智能,作者|魏琳華,編輯|劉俊宏AI圈的規則就是用來打破的,市場變化之快,讓公司們經常猝不及防。

DeepSeek 傳將向聯合國安理會閉門通報 AI 潛在風險與安全挑戰
據路透社援引知情人士消息報道,中國人工智能初創企業 DeepSeek 預計將於本週向聯合國安理會就人工智能所帶來的潛在風險與安全挑戰進行專題通報。作為全球矚目的前沿 AI 力量,此次 DeepSeek 受邀參與聯合國安理會的安全通報,折射出國際社會對中國大模型企業在技術治理、風險防控以及全球 AI 安全框架建設中發揮實質性作用的重視。

Yann LeCun 萬字演講:「預測像素」是偽命題,JEPA 也並非憑空而來 | ECCV 2026
本文作者: 叢末 2026-09-23 09:49 導語:只靠語言和 token,AI 永遠跨不過物理世界這道門檻!只靠語言和 token,AI 永遠跨不過物理世界這道門檻!作者丨幸麗娟 編輯丨岑 峰 AI 行業的主敘事,曾被 Scaling Law 壟斷:更大的模型、更多的 token,更強的文本推理,彷彿只要把語言模型繼續堆大,通用人工智能就會自動到來。