Anthropic 發布 Claude Sonnet 5.5:Terminal-Bench 4.0 得分 70.6%,維持 $2/$10 定價
Anthropic just released Claude Sonnet 5.5.It is the second model in the Claude 5.5 family, following Claude Opus 5.5.Anthropic positions it as a faster, lower-cost complement to Opus 5.5.It targets well-scoped everyday tasks, bug fixing, and polished documents, slides, and spreadsheets.
Is it deployable?Yes, It is live on the Claude Platform as claude-sonnet-5-5, plus AWS, Google Cloud, and Microsoft Azure.It is a closed-weights model, so self-hosting is not an option.
What Changed Versus Sonnet 5 Anthropic reports 4 main upgrades over Sonnet 5: Speed: output generation is 30%+ faster, making it the fastest Sonnet to date.Cost per task: up to 30% lower, because it needs fewer tokens and tool calls.
Writing: clearer prose, with early testers calling it a better collaboration partner.Vision and long-horizon work: it is the first Sonnet to beat Pokémon Red using only screenshots.Specs from the models overview: 1M-token context, 128K max output, and a June 2026 reliable knowledge cutoff.
Adaptive thinking is on by default.Effort runs across 5 levels: low, medium, high, xhigh, and max.Benchmarks All scores below are vendor-reported in the launch post.Methodology lives in the Sonnet 5.5 System Card.Terminal-Bench 4.0: 70.6%, versus 10.3% for Sonnet 5 and 66.4% for Opus 5.5 at Xhigh.
CursorBench 4.0: 55.5%, about 2 points below Opus 5.5 (57.8%).FrontierCode 1.1: 52.1% at Xhigh and 46.2% at Max.GPT-6 Sol scored 49.3%.GDPval-AA v2.1: 1844, versus 1846 for Opus 5.5 and 1449 for Sonnet 5.OSWorld 2.1: 80.1% on computer use, close to Opus 5.5 (81.8%).Humanity’s Last Exam: 64.
5% with tools, up from 54.9%.The Max result is lower than Xhigh for a reason.At Max, the model more often ran multi-agent code review.That sometimes caused timeouts or out-of-scope edits, which FrontierCode penalizes.Anthropic also states that Opus 5.
5 remains clearly stronger on complex, open-ended work.Pricing and Efficiency List pricing is unchanged from Sonnet 5: $2 per million input tokens and $10 per million output tokens.Cache reads cost $0.20 and cache writes $2.50 per million.That is half of Opus 5.5 ($4/$20).
Savings come from token efficiency, not a price cut.Customer data supports this.Balyasny Asset Management measured about 121K tokens per answer versus 497K on Sonnet 5.Base44 reported 3.6 iterations per app build, where Opus 5 took 7.7.Zendesk processed tickets 20% faster.
Effort defaults differ by surface.Claude Code and the Claude apps default to Medium.The Claude Platform defaults to High.How It Compares FeatureClaude Sonnet 5.5GPT-6 SolGemini 3.1 Pro PreviewClaude Opus 5.
5DeveloperAnthropicOpenAIGoogleAnthropicPrice per 1M tokens (input/output)$2 / $10$2 / $10 (prompts up to 272K)$2 / $12 (prompts up to 200K)$4 / $20Context window1M tokens1,050,000 tokens1,048,576 tokens1M tokensMax output128K tokens128K tokens65,536 tokens128K tokensKnowledge cutoffJune 2026April 20, 2026Not listedJune 2026Reasoning controlAdaptive thinking, 5 effort levels6 effort levels (none to max)Thinking supportedAdaptive thinking (always on)InputsText, imageText, imageText, image, video, audio, PDFText, imageFrontierCode 1.
152.1% (Xhigh)49.3%Not reported54.4%GDPval-AA v2.118441487Not reported1846Chartography (no tools)61.6%53.6%Not reported64.
4%Release stageGenerally availableGenerally availablePreviewGenerally availableOpen weightsNoNoNoNo Benchmark scores are vendor-reported by Anthropic; GDPval-AA runs by Artificial Analysis.Prices are standard API list rates, verified September 28, 2026.
Interactive Explainer #mtp-s55{background:#141413!important;color:#FAF9F5!important;font-family:ui-sans-serif,system-ui,-apple-system,"Segoe UI",Roboto,Arial,sans-serif;border:1px solid #2b2a27!
important;border-radius:14px;padding:22px;max-width:860px;margin:0 auto;box-sizing:border-box} #mtp-s55 *{box-sizing:border-box} #mtp-s55 .hd{display:flex;align-items:center;gap:10px;margin-bottom:4px} #mtp-s55 .
dot{width:12px;height:12px;border-radius:50%;background:#D97757;animation:pulse 2s infinite} @keyframes pulse{0%{box-shadow:0 0 0 0 rgba(217,119,87,.
6)}70%{box-shadow:0 0 0 10px rgba(217,119,87,0)}100%{box-shadow:0 0 0 0 rgba(217,119,87,0)}} #mtp-s55 h3{margin:0;font-size:20px;font-family:Georgia,"Times New Roman",serif;font-weight:500;color:#FAF9F5!important} #mtp-s55 .sub{color:#B0AEA5;font-size:13px;margin:0 0 16px} #mtp-s55 .
tabs{display:flex;flex-wrap:wrap;gap:6px;margin-bottom:16px} #mtp-s55 .tab{background:#1f1e1b!important;color:#B0AEA5!important;border:1px solid #2f2e2a!important;border-radius:999px;padding:8px 14px;font-size:13px;cursor:pointer;transition:all .2s} #mtp-s55 .tab:hover{color:#FAF9F5!
important;border-color:#D97757!important} #mtp-s55 .tab.on{background:#D97757!important;color:#141413!important;border-color:#D97757!important;font-weight:600} #mtp-s55 .pane{display:none;animation:fade .35s ease} #mtp-s55 .pane.
on{display:block} @keyframes fade{from{opacity:0;transform:translateY(6px)}to{opacity:1;transform:none}} #mtp-s55 .row{display:grid;grid-template-columns:110px 1fr 64px;align-items:center;gap:10px;margin:9px 0;font-size:13px} #mtp-s55 .trk{background:#26251f!
important;height:14px;border-radius:7px;overflow:hidden} #mtp-s55 .bar{height:100%;width:0;border-radius:7px;transition:width .9s cubic-bezier(.2,.8,.2,1)} #mtp-s55 .val{text-align:right;font-variant-numeric:tabular-nums;color:#FAF9F5} #mtp-s55 .
chips{display:flex;flex-wrap:wrap;gap:6px;margin-bottom:10px} #mtp-s55 .chip{background:transparent!important;color:#B0AEA5!important;border:1px solid #3a3934!important;border-radius:8px;padding:6px 10px;font-size:12px;cursor:pointer} #mtp-s55 .chip.on{border-color:#6A9BCC!important;color:#FAF9F5!
important;background:#1b2530!important} #mtp-s55 .note{font-size:12px;color:#8f8d86;margin-top:10px;line-height:1.5} #mtp-s55 .card{background:#1c1b18!important;border:1px solid #2b2a27!
important;border-radius:10px;padding:14px;margin-top:10px} #mtp-s55 label{font-size:13px;color:#B0AEA5;display:block;margin:10px 0 4px} #mtp-s55 input[type=range]{width:100%;accent-color:#D97757} #mtp-s55 .big{font-size:26px;font-weight:600;color:#D97757;font-variant-numeric:tabular-nums} #mtp-s55 .
grid2{display:grid;grid-template-columns:1fr 1fr;gap:10px} #mtp-s55 .k{font-size:12px;color:#B0AEA5} #mtp-s55 .dial{display:flex;gap:6px;margin:8px 0} #mtp-s55 .lv{flex:1;text-align:center;padding:10px 4px;border-radius:8px;background:#1f1e1b!important;border:1px solid #2f2e2a!
important;font-size:12px;cursor:pointer;color:#B0AEA5!important;transition:all .25s} #mtp-s55 .lv.on{background:#D97757!important;color:#141413!important;font-weight:600;transform:translateY(-2px)} #mtp-s55 .meter{height:10px;border-radius:5px;background:#26251f!
important;overflow:hidden;margin:4px 0 10px} #mtp-s55 .mfill{height:100%;transition:width .6s ease} #mtp-s55 pre,#mtp-s55 code{background:#0e0e0d!important;color:#E8E6DC!important;border:1px solid #2b2a27!important;border-radius:8px;font-family:ui-monospace,Menlo,Consolas,monospace!
important;font-size:12px!important} #mtp-s55 pre{padding:10px;margin:6px 0;white-space:pre-wrap;word-break:break-word} #mtp-s55 .bad{color:#E07A6B!important} #mtp-s55 .ok{color:#9DB57A!important} #mtp-s55 .ft{margin-top:18px;padding-top:10px;border-top:1px solid #2b2a27!
important;font-size:12px;color:#8f8d86;display:flex;justify-content:space-between;flex-wrap:wrap;gap:6px} #mtp-s55 .ft b{color:#D97757} #mtp-s55 hr,#mtp-s55 p:empty,#mtp-s55 del,#mtp-s55 s{display:none!important} @media (max-width:640px){#mtp-s55{padding:14px}#mtp-s55 .
row{grid-template-columns:84px 1fr 54px;font-size:12px}#mtp-s55 .grid2{grid-template-columns:1fr}#mtp-s55 .lv{font-size:11px;padding:8px 2px}} Claude Sonnet 5.5, explained interactivelyTap through benchmarks, effort levels, task cost and the API migration checker.
BenchmarksEffort dialCost per taskMigration checkerSelect an effort level.Sonnet 5.5 exposes 5.Reasoning depthSpeed and token savingsMeters are illustrative of the direction Anthropic descri
Related
相關文章

IDC:今年上半年全球人形機器人出貨近 2.5 萬臺同比增長 432.1%,中國獨佔 1.9 萬臺
作者:小泵 責編:小泵 評論: 9 月 29 日消息,國際數據公司(IDC)最新發布的全球人形機器人跟蹤報告顯示,2026 上半年,全球人形機器人出貨量接近 2.5 萬臺,同比增長 432.1%,產業發展進一步從技術驗證向商業化應用探索階段邁進。

微軟報告:全球打工人裡每五人就有一個用 AI,南北差距還在拉大
數據顯示,到今年第二季度,全世界勞動年齡人口中已有 18.8% 用上了人工智能,比上一季度多出約 1 個百分點,且幾乎每個經濟體的使用率都在往上走。阿聯酋、新加坡領跑,日韓追趕最快把鏡頭拉到地區層面,最捨得用 AI 的還是阿聯酋(70.1%)和新加坡(64.

微軟納德拉:AI 行業“有點飄了”,呼籲迴歸用戶需求
作者:故淵 責編:故淵 評論: 9 月 29 日消息,亞歷克斯 · 希思(Alex Heath)於 9 月 27 日在 LinkedIn 平臺發佈視頻,分享了一段採訪微軟首席執行官薩蒂亞 · 納德拉(Satya Nadella)的視頻,聚焦當前行業 AI 發展現狀以及未來趨勢等。

Fireworks AI 推出 FireRouter with Opus:編碼任務成本降 57%,準確率僅損失 1.5 個百分點
經過一個多月的內部 A/B 測試,該方案執行編碼任務的準確率達到單獨使用 Opus 的98.1%,而成本降低了57%。FireRouter 的核心邏輯是在每次用戶輪次中,評估模型集合中每個模型對當前任務的適合程度,估算每個模型的處理成本(包括提示緩存成本),並權衡切換模型帶來的節省與丟失緩存是否值得,最終路由到質量與成本之間取得最佳平衡的模型。

Anthropic搶在OpenAI開發者大會前甩出Claude Sonnet 5.5,一半價格逼近Opus、還通關了寶可夢
這款被定位為日常工作全能助手的新模型,把速度快、成本低、質量高這套常被說成不可能三角的標籤直接撕了下來——相比前代 Sonnet5,速度提升30%,單任務成本最高降30%,定價卻只有 Opus5.5的一半,性能卻幾乎追平,甚至在 Terminal-Bench4.

AI 寫作特徵減少:Claude Opus 5.5 破折號使用量下降約 95%、分號下降 73%
作者:故淵 責編:故淵 評論: 9 月 29 日消息,@arena 於 9 月 26 日在 X 平臺發佈推文,通過寫作指標測試,發現相比較 Opus 5 模型,Anthropic 的 Claude Opus 5.5 破折號使用量下降約 95%,每 1,000 個單詞從 15.