Z.ai 推出 GLM-5.3,無需重新訓練基礎模型:複雜程式編寫與長期任務表現更佳
Z.ai just released GLM-5.3.GLM-5.3 runs on the same 743B base model as GLM-5.2.Every reported gain comes from scaled post-training: more task environments, more environment types, longer training.The results land in two places.
Coding jumps most on the longest-horizon benchmarks, with Terminal-Bench 3.0 moving from 4.6 to 28.3.Cybersecurity moved further than Z.ai says it expected, with CyberGym reaching 84.5%.Weights are not public yet.Is It Deployable?Partially, GLM-5.3 is live through the Z.
ai API, the GLM Coding Plan, and ZCode.Weights are not out.Z.ai says it will publish them roughly two weeks after launch, once safety evaluation and hardening finish.Which companies can move now: Startups and mid-market engineering orgs can adopt it today via the Coding Plan or API.
Enterprises with data-residency or vendor-review rules should wait for weights.Security vendors and MSSPs get the most signal, and the most policy exposure.
Industries: Developer tooling, cloud infrastructure, application security, fintech and e-commerce engineering, and vendors shipping kernels, browser engines, or network stacks.
Applications: Repository-scale refactors, long-horizon CLI agents, CI failure triage, white-box vulnerability discovery, crash triage, and secure code review.Coding Results Terminal-Bench 3.0 moves from 4.6 to 28.3 against GLM-5.2.DeepSWE v1.1 moves from 46.2 to 66.9.
Agents’ Last Exam (CLI) moves from 23.8 to 28.5.On GDPval-AA v2, which spans 44 occupations, GLM-5.3 scores 1,769.On Z.ai Code Bench, an internal evaluation, the company reports a 50% improvement over GLM-5.2.It reports 31.4% at roughly 50,000 output tokens per task.Claude Opus 4.8 scores 29.
5% at 120,000 tokens.Claude Fable 5 still leads at 39.5% at maximum effort.Z.ai argues a private benchmark reduces contamination risk.On public suites, GLM-5.3 trails GPT-5.6 Sol and Fable 5 on several harder coding evaluations.
All figures are vendor-reported, with harness, context length, and sampling settings documented in the announcement.The Cybersecurity Result Z.ai flags this one as unplanned.It added vulnerability-discovery data expecting better single-bug reasoning.
Instead, capability kept compounding as training scaled.The model began forming coherent plans across complete exploitation chains.CyberGym, which tests discovery and validation from white-box source, moves from 77.2% to 84.5%.That edges past Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%.
ExploitBench, which requires root-cause reasoning and a working exploit, moves from 24.4% to 54.4%.Mythos 5 sits at 78.0%.On ExploitGym, GLM-5.3 completes 105 tasks in two hours and 130 in six.GLM-5.2 completes 29 and 39.Mythos 5 completes 181 and 247.The pattern is consistent.
The deeper into the exploitation chain a benchmark sits, the larger the gain over GLM-5.2.The gap to closed frontier models also widens.Interactive Explainer (function(){var f=document.getElementById("mtp-glm53"); window.addEventListener("message",function(e){var d=e.data; if(d&&d.
mtpEmbed==="glm53"&&typeof d.height==="number"&&d.height>200&&d.height<6000){f.style.height=d.height+"px";}}); })(); Key Takeaways GLM-5.3 reuses the GLM-5.2 base model; all gains come from post-training scaling.Terminal-Bench 3.0 moves from 4.6 to 28.3; DeepSWE v1.1 from 46.2 to 66.9.
CyberGym hits 84.5%, ahead of Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%).ExploitBench more than doubles to 54.4%, but trails Mythos 5 at 78.0%.Weights ship in about two weeks, after safety evaluation and hardening.Check out the Z.ai GLM-5.3 technical blog, Zai_org announcement, Z.
ai Security Disclosure Ledger and zai-org/GLM-5 on GitHub.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?Connect with us The post Z.ai Ships GLM-5.3 Without Retraining the Base Model: Better at Complex Coding and Long-Horizon Tasks appeared first on MarkTechPost.
Related
相關文章

2026 世界機器人大會在京閉幕:首發新品 311 件,下一屆官宣 2027 年 8 月 18 日開幕
作者:汪淼 責編:汪淼 評論: 8 月 24 日消息,為期 5 天的 2026 世界機器人大會於 8 月 23 日在北京落下帷幕。

買場景、建團隊、借渠道,OpenAI如何做企業服務
17分鐘前為什麼海外市場成了影響泡泡瑪特價值的關鍵18分鐘前從高珠到大眾時尚,飾品行業迎來新變化2026-08-20閱讀更多內容,狠戳這裡查看AI測評豆包WorkBuddy千問選靠譜AI,看真實評測查看AI測評官方交流社區加入諮詢項目審核和入駐聯繫項目推薦訂閱號關注下一篇格局生變,新銳集體圍剿韓妝巨頭?

李澤湘投過的商業園林機器人完成數千萬融資,瞄準海外綠地智能運維|硬氪首發 |
李澤湘投資的商業園林機器人公司完成數千萬人民幣融資,目標鎖定海外綠地智能運維市場。該公司提供商業園林智能作業平台,強調可複製且能降低成本、提升效率的生產力方案。
曝Hugging Face擬出售,估值或達130億美元
Hugging Face尚未對出售消息作出正式回應。Hugging Face是全球最大的開源大語言模型和數據集託管平臺,被業內視為“AI領域的GitHub”。若此次以130億美元或更高的估值完成出售,Hugging Face的估值將在不到三年的時間內增長近三倍,反映出AI基礎設施平臺在行業中的戰略價值正在快速攀升。就在此次出售消息傳出前一個月,7月16日Hugging Face剛剛經歷了一場震驚業界的安全事件。OpenAI的AI模型在安全測試中“失控”,自主入侵了Hugging Face的生產基礎設施。

脫單也靠AI,“賽博媒婆”拿到1500萬元天使融資
吳維消費星球2026.08.24 12:08 · 來自四川全文5561字00:00 / 16:05AI是婚戀的良配嗎?文 | 吳維消費星球2026年的創投圈,有一條心照不宣的潛規則:BP裡寫上”AI”兩個字,融資路能少走一半。AI賣咖啡,AI看風水,AI算八字,AI陪失眠的年輕人聊天到天亮。
英偉達AI服務器曝將漲價超15% 內存成本飆升成核心推手
據彭博社消息,英偉達 已告知其部分最大客戶,搭載AI芯片的服務器價格將普遍上漲超過15%,主要原因是內存芯片成本急劇飆升。具體漲價幅度將取決於英偉達芯片的型號以及內存配置。英偉達方面對此尚未作出回應。據分析師預測,2026年第二季度傳統DRAM合約價格環比將上漲58%至63%,此前第一季度已飆升90%至95%。消息面上,英偉達將於8月26日公佈第二財季財報。業內人士認為,此次漲價將進一步推高AI數據中心的建設成本。