Anthropic 發布 Claude Sonnet 5.5:Terminal-Bench 4.0 得分 70.6%,維持 $2/$10 定價
Anthropic just released Claude Sonnet 5.5.It is the second model in the Claude 5.5 family, following Claude Opus 5.5.Anthropic positions it as a faster, lower-cost complement to Opus 5.5.It targets well-scoped everyday tasks, bug fixing, and polished documents, slides, and spreadsheets.
Is it deployable?Yes, It is live on the Claude Platform as claude-sonnet-5-5, plus AWS, Google Cloud, and Microsoft Azure.It is a closed-weights model, so self-hosting is not an option.
What Changed Versus Sonnet 5 Anthropic reports 4 main upgrades over Sonnet 5: Speed: output generation is 30%+ faster, making it the fastest Sonnet to date.Cost per task: up to 30% lower, because it needs fewer tokens and tool calls.
Writing: clearer prose, with early testers calling it a better collaboration partner.Vision and long-horizon work: it is the first Sonnet to beat Pokémon Red using only screenshots.Specs from the models overview: 1M-token context, 128K max output, and a June 2026 reliable knowledge cutoff.
Adaptive thinking is on by default.Effort runs across 5 levels: low, medium, high, xhigh, and max.Benchmarks All scores below are vendor-reported in the launch post.Methodology lives in the Sonnet 5.5 System Card.Terminal-Bench 4.0: 70.6%, versus 10.3% for Sonnet 5 and 66.4% for Opus 5.5 at Xhigh.
CursorBench 4.0: 55.5%, about 2 points below Opus 5.5 (57.8%).FrontierCode 1.1: 52.1% at Xhigh and 46.2% at Max.GPT-6 Sol scored 49.3%.GDPval-AA v2.1: 1844, versus 1846 for Opus 5.5 and 1449 for Sonnet 5.OSWorld 2.1: 80.1% on computer use, close to Opus 5.5 (81.8%).Humanity’s Last Exam: 64.
5% with tools, up from 54.9%.The Max result is lower than Xhigh for a reason.At Max, the model more often ran multi-agent code review.That sometimes caused timeouts or out-of-scope edits, which FrontierCode penalizes.Anthropic also states that Opus 5.
5 remains clearly stronger on complex, open-ended work.Pricing and Efficiency List pricing is unchanged from Sonnet 5: $2 per million input tokens and $10 per million output tokens.Cache reads cost $0.20 and cache writes $2.50 per million.That is half of Opus 5.5 ($4/$20).
Savings come from token efficiency, not a price cut.Customer data supports this.Balyasny Asset Management measured about 121K tokens per answer versus 497K on Sonnet 5.Base44 reported 3.6 iterations per app build, where Opus 5 took 7.7.Zendesk processed tickets 20% faster.
Effort defaults differ by surface.Claude Code and the Claude apps default to Medium.The Claude Platform defaults to High.How It Compares FeatureClaude Sonnet 5.5GPT-6 SolGemini 3.1 Pro PreviewClaude Opus 5.
5DeveloperAnthropicOpenAIGoogleAnthropicPrice per 1M tokens (input/output)$2 / $10$2 / $10 (prompts up to 272K)$2 / $12 (prompts up to 200K)$4 / $20Context window1M tokens1,050,000 tokens1,048,576 tokens1M tokensMax output128K tokens128K tokens65,536 tokens128K tokensKnowledge cutoffJune 2026April 20, 2026Not listedJune 2026Reasoning controlAdaptive thinking, 5 effort levels6 effort levels (none to max)Thinking supportedAdaptive thinking (always on)InputsText, imageText, imageText, image, video, audio, PDFText, imageFrontierCode 1.
152.1% (Xhigh)49.3%Not reported54.4%GDPval-AA v2.118441487Not reported1846Chartography (no tools)61.6%53.6%Not reported64.
4%Release stageGenerally availableGenerally availablePreviewGenerally availableOpen weightsNoNoNoNo Benchmark scores are vendor-reported by Anthropic; GDPval-AA runs by Artificial Analysis.Prices are standard API list rates, verified September 28, 2026.
Interactive Explainer #mtp-s55{background:#141413!important;color:#FAF9F5!important;font-family:ui-sans-serif,system-ui,-apple-system,"Segoe UI",Roboto,Arial,sans-serif;border:1px solid #2b2a27!
important;border-radius:14px;padding:22px;max-width:860px;margin:0 auto;box-sizing:border-box} #mtp-s55 *{box-sizing:border-box} #mtp-s55 .hd{display:flex;align-items:center;gap:10px;margin-bottom:4px} #mtp-s55 .
dot{width:12px;height:12px;border-radius:50%;background:#D97757;animation:pulse 2s infinite} @keyframes pulse{0%{box-shadow:0 0 0 0 rgba(217,119,87,.
6)}70%{box-shadow:0 0 0 10px rgba(217,119,87,0)}100%{box-shadow:0 0 0 0 rgba(217,119,87,0)}} #mtp-s55 h3{margin:0;font-size:20px;font-family:Georgia,"Times New Roman",serif;font-weight:500;color:#FAF9F5!important} #mtp-s55 .sub{color:#B0AEA5;font-size:13px;margin:0 0 16px} #mtp-s55 .
tabs{display:flex;flex-wrap:wrap;gap:6px;margin-bottom:16px} #mtp-s55 .tab{background:#1f1e1b!important;color:#B0AEA5!important;border:1px solid #2f2e2a!important;border-radius:999px;padding:8px 14px;font-size:13px;cursor:pointer;transition:all .2s} #mtp-s55 .tab:hover{color:#FAF9F5!
important;border-color:#D97757!important} #mtp-s55 .tab.on{background:#D97757!important;color:#141413!important;border-color:#D97757!important;font-weight:600} #mtp-s55 .pane{display:none;animation:fade .35s ease} #mtp-s55 .pane.
on{display:block} @keyframes fade{from{opacity:0;transform:translateY(6px)}to{opacity:1;transform:none}} #mtp-s55 .row{display:grid;grid-template-columns:110px 1fr 64px;align-items:center;gap:10px;margin:9px 0;font-size:13px} #mtp-s55 .trk{background:#26251f!
important;height:14px;border-radius:7px;overflow:hidden} #mtp-s55 .bar{height:100%;width:0;border-radius:7px;transition:width .9s cubic-bezier(.2,.8,.2,1)} #mtp-s55 .val{text-align:right;font-variant-numeric:tabular-nums;color:#FAF9F5} #mtp-s55 .
chips{display:flex;flex-wrap:wrap;gap:6px;margin-bottom:10px} #mtp-s55 .chip{background:transparent!important;color:#B0AEA5!important;border:1px solid #3a3934!important;border-radius:8px;padding:6px 10px;font-size:12px;cursor:pointer} #mtp-s55 .chip.on{border-color:#6A9BCC!important;color:#FAF9F5!
important;background:#1b2530!important} #mtp-s55 .note{font-size:12px;color:#8f8d86;margin-top:10px;line-height:1.5} #mtp-s55 .card{background:#1c1b18!important;border:1px solid #2b2a27!
important;border-radius:10px;padding:14px;margin-top:10px} #mtp-s55 label{font-size:13px;color:#B0AEA5;display:block;margin:10px 0 4px} #mtp-s55 input[type=range]{width:100%;accent-color:#D97757} #mtp-s55 .big{font-size:26px;font-weight:600;color:#D97757;font-variant-numeric:tabular-nums} #mtp-s55 .
grid2{display:grid;grid-template-columns:1fr 1fr;gap:10px} #mtp-s55 .k{font-size:12px;color:#B0AEA5} #mtp-s55 .dial{display:flex;gap:6px;margin:8px 0} #mtp-s55 .lv{flex:1;text-align:center;padding:10px 4px;border-radius:8px;background:#1f1e1b!important;border:1px solid #2f2e2a!
important;font-size:12px;cursor:pointer;color:#B0AEA5!important;transition:all .25s} #mtp-s55 .lv.on{background:#D97757!important;color:#141413!important;font-weight:600;transform:translateY(-2px)} #mtp-s55 .meter{height:10px;border-radius:5px;background:#26251f!
important;overflow:hidden;margin:4px 0 10px} #mtp-s55 .mfill{height:100%;transition:width .6s ease} #mtp-s55 pre,#mtp-s55 code{background:#0e0e0d!important;color:#E8E6DC!important;border:1px solid #2b2a27!important;border-radius:8px;font-family:ui-monospace,Menlo,Consolas,monospace!
important;font-size:12px!important} #mtp-s55 pre{padding:10px;margin:6px 0;white-space:pre-wrap;word-break:break-word} #mtp-s55 .bad{color:#E07A6B!important} #mtp-s55 .ok{color:#9DB57A!important} #mtp-s55 .ft{margin-top:18px;padding-top:10px;border-top:1px solid #2b2a27!
important;font-size:12px;color:#8f8d86;display:flex;justify-content:space-between;flex-wrap:wrap;gap:6px} #mtp-s55 .ft b{color:#D97757} #mtp-s55 hr,#mtp-s55 p:empty,#mtp-s55 del,#mtp-s55 s{display:none!important} @media (max-width:640px){#mtp-s55{padding:14px}#mtp-s55 .
row{grid-template-columns:84px 1fr 54px;font-size:12px}#mtp-s55 .grid2{grid-template-columns:1fr}#mtp-s55 .lv{font-size:11px;padding:8px 2px}} Claude Sonnet 5.5, explained interactivelyTap through benchmarks, effort levels, task cost and the API migration checker.
BenchmarksEffort dialCost per taskMigration checkerSelect an effort level.Sonnet 5.5 exposes 5.Reasoning depthSpeed and token savingsMeters are illustrative of the direction Anthropic descri
Related
相關文章

字節豆包要出獨立個人助理App了,內部已秘密內測數月
AI資訊AI新聞資訊正文字節豆包要出獨立個人助理App了,內部已秘密內測數月發佈於AI新聞資訊發佈時間 :2026年9月29號 14:14閱讀 :1分鐘字節跳動旗下的AI產品豆包,正加速向獨立個人助理方向邁進。據知情人士透露,豆包從今年4月起就已開始秘密內測一款名為“Spell”的個人助理應用,近期隨著海外同類產品熱度攀升,字節內部明顯加快了這一項目的推進節奏。據瞭解,字節計劃將原有的“Spell”項目團隊與豆包對話團隊進行合併,集中力量打造一款獨立的個人助理App。這意味著豆包不再只是內嵌在其它產品中的對話工具,而是要以獨立應用形態直接面向用戶。目前,該產品尚未公佈正式上線時間,具體功能細節也仍處於保密階段。不過從字節的動作來看,AI個人助理賽道正成為各大廠商爭奪的下一個焦點。相關推薦微軟報告:全球打工人裡每五人就有一個用 AI,南北差距還在拉大微軟21日發佈《全球人工智能普及報告》稱,今年二季度全球勞動年齡人口中18.8%已使用人工智能,較上季度增約1個百分點,幾乎所有經濟體使用率都在上升。阿聯酋(70.1%)和新加坡(64.3%)領跑,愛爾蘭、法國、挪威接近一半;日韓追趕最快,其中韓國增速最猛。2026年9月29號 11:06136.8kAnthropic搶在OpenAI開發者大會前甩出Claude Sonnet 5.5,一半價格逼近Opus、還通關了寶可夢Anthropic 搶在 OpenAI 年度開發者大會前發佈 Claude Sonnet5.5,定位日常工作全能助手。它打破速度、成本、質量“不可能三角”:較 Sonnet5 提速30%,單任務成本最高降30%,定價僅 Opus5.5 一半,性能幾乎追平,並在 Terminal-Bench4.0 等多項測試中反超自家高端型號;文章還以遊戲為例說明其“腦子夠用”。2026年9月29號 10:38148.8kKling

我國生成式人工智能用戶規模突破 7 億人,普及率超 50%
首頁 > 智能時代>人工智能 我國生成式人工智能用戶規模突破 7 億人,普及率超 50% 2026/9/29 14:16:03 來源:IT之家 作者:遠洋 責編:遠洋 評論: 感謝IT之家網友 HH_KK、很宅很怕生、小星_14 的線索投遞! IT之家 9 月 29 日消息,據央視新聞報道,9 月 29 日,在 2026(第七屆)中國互聯網基礎資源大會上,中國互聯網絡信息中心政策與國際合作所發佈《生成式人工智能應用發展報告(2026)》。報告顯示,截至 2026 年上半年,我國生成式人工智能用戶規模已超過 7 億人,普及率突破 50%。報告顯示,智能問答仍是生成式人工智能最主流的使用方式,76% 的用戶會藉助其獲取答案;圖片和視頻處理、文本處理、工作總結及會議紀要生成等應用場景的用戶佔比也分別達到 47.8%、37.6% 和 32.5%。IT之家注意到,從使用頻次看,AI 綜合助手、AI 效率辦公類應用的使用次數同比均實現翻倍以上增長,反映出用戶對智能化工具的依賴程度持續加深。智能消費方面,38.7% 的網民近半年內網購過智能硬件設備。其中,可穿戴設備和 3C 數碼產品是我國網民接觸智能硬件設備的首要入口。購買過智能可穿戴設備的網民比例為 20.2%;購買智能手機、平板電腦等 3C 數碼產品的網民比例為 18.2%。報告指出,我國在人工智能全產業鏈多個關鍵環節實現快速突破。算力基礎設施方面,截至 2026 年 6 月,我國智能算力規模達到 2185EFLOPS,同比大幅增長 177%。尤為值得關注的是,首個全國產 10 萬卡人工智能超集群已正式投入運行,其芯片、計算、存儲、網絡、散熱等核心軟硬件全部採用國產自主可控技術,實現從底層硬件到上層服務的全棧國產化,我國算力基礎設施建設由此邁入十萬卡級部署的新階段。模型能力方面,深度求索、月之暗面等本土企業相繼推出多個萬億參數級開源

AI手機真正要消滅的,是“操作手機”這件事
防冷塗的蠟2026.09.29 13:14 · 來自浙江全文5663字00:00 / 16:00一個已經能聽懂“我要什麼”的機器,為什麼還要學習“下一步點哪裡”?文 | 防冷塗的蠟現在的AI手機,正在變得越來越會“辦事”。對著手機說一句“幫我點一杯熱拿鐵”,然後看它自己打開外賣App、找到常去的咖啡店、完成選購,最後等你確認付款。

IDC:今年上半年全球人形機器人出貨近 2.5 萬臺同比增長 432.1%,中國獨佔 1.9 萬臺
作者:小泵 責編:小泵 評論: 9 月 29 日消息,國際數據公司(IDC)最新發布的全球人形機器人跟蹤報告顯示,2026 上半年,全球人形機器人出貨量接近 2.5 萬臺,同比增長 432.1%,產業發展進一步從技術驗證向商業化應用探索階段邁進。

微軟報告:全球打工人裡每五人就有一個用 AI,南北差距還在拉大
數據顯示,到今年第二季度,全世界勞動年齡人口中已有 18.8% 用上了人工智能,比上一季度多出約 1 個百分點,且幾乎每個經濟體的使用率都在往上走。阿聯酋、新加坡領跑,日韓追趕最快把鏡頭拉到地區層面,最捨得用 AI 的還是阿聯酋(70.1%)和新加坡(64.

微軟納德拉:AI 行業“有點飄了”,呼籲迴歸用戶需求
作者:故淵 責編:故淵 評論: 9 月 29 日消息,亞歷克斯 · 希思(Alex Heath)於 9 月 27 日在 LinkedIn 平臺發佈視頻,分享了一段採訪微軟首席執行官薩蒂亞 · 納德拉(Satya Nadella)的視頻,聚焦當前行業 AI 發展現狀以及未來趨勢等。