Claude自動對齊進展
AlignmentAutomated researchers can reliably mitigate alignment failuresAug 28, 2026Read the paperAs AI begins to build itself, automating alignment research becomes increasingly important to let safety research keep pace.
Although measuring the success of alignment research is enormously challenging, researchers (at Anthropic and elsewhere) have developed benchmarks and automated auditing tools, such as Petri, that quantify common alignment failures, like deception, sycophancy, and jailbreaks.
In one of our earlier experiments, we tasked Claude with finding effective ways to use weak AI models as “teachers” to supervise the training of stronger models (in this case, the “student” model).Now, we’re releasing a new report that builds on this idea.
We had Claude autonomously train models to improve their performance on several public benchmarks that measure each of 10 categories of alignment failure.For instance, Claude improved models’ performance on privacy violation, measured by ConfAIde, PrivaCI-Bench, and PrivacyLens.
Claude tackled one alignment failure at a time through a loop of searching literature, proposing methods and data, training, and then testing.We judged Claude’s success according to the “percentage of safety gap closed,” i.e.
, how far its methods moved the student model towards the theoretical perfect score, as judged across the range of benchmarks (typically three to five) for each category of alignment failure.
We excluded alignment methods that hurt the student models’ general capabilities, and forbade Claude from distilling its own alignment directly into the target model.We enforced these constraints with a monitoring agent, which read every method Claude had in mind before it ran.
Our aim was to assess whether the proposed methods would, first, remain effective on alignment evaluations that Claude was never shown during its research loop; second, avoid degrading the student model’s capabilities (since safety training might, for example, make models refuse tasks more often, reducing their overall usability); and, third, still work on larger models than the ones Claude was asked to align in this test.
On each of these counts, Claude’s methods worked.For all 10 alignment failures, Claude found fixes that improved the target benchmarks without degrading capabilities.
The best methods also worked on withheld alignment benchmarks and on Petri, an open-source tool that simulates adversarial multi-turn scenarios for testing misalignment.Moreover, the methods remained effective on models up to 4.7 times larger than those Claude optimized for during the research loop.
Successfully mitigating diverse alignment failures.We applied the automated alignment researcher to mitigate 10 alignment failures separately, and in each scenario it closed a substantial portion of the safety gap to perfect performance.
Claude also outscored 28 human safety researchers who had up to eight hours to devise methods.On deception, for example, Claude’s best method performed 20% better than the best human proposal.
However, since the humans couldn’t iterate on their submissions, we view this less as a direct comparison and more as evidence for a workflow where Claude identifies promising alignment methods that humans can refine further.Automated alignment researchers mitigating deception in Gemma-2-2B.
Claude submitted more than 150 attempts at mitigating deceptive behavior, and achieved a final performance of 82% of the safety gap closed in this run.On average, it achieved 85% across multiple runs.
In contrast, six experienced safety researchers working under the same rules proposed methods that closed 20% of the gap to a perfect score, on average, on the benchmarks the methods were trained against.(Error bars are 95% confidence intervals.
)In the future, when Claude becomes better at alignment research than even the best human researchers, we might want Claude to directly align its stronger successors.To assess this, we evaluated whether a weaker Claude model could mitigate alignment failures in more powerful ones.
Can Claude post-train a production-grade model for better alignment?We tasked Claude Sonnet 5—which is weaker than Claude Opus 4.8 on the Epoch Capabilities Index, a metric that considers comprehensive capability dimensions—with fixing alignment failures in an early Opus 4.
8 checkpoint that had not yet gone through most of our production alignment training.In just 60 hours, Claude experimented with over 50 solutions and achieved alignment scores nearly matching those of our production models.
The winning solution contains just over 2,000 training examples, built from simple templates or public datasets, making it roughly 15,000 times more efficient than our production alignment procedure.Applying the automated alignment researcher to a production system.
When Claude Sonnet 5 proposed mitigations for alignment failures in an early checkpoint of Claude Opus 4.8, it quickly discovered a method that closed most of the safety gap with the released version of Opus 4.8.
What Claude proposedFor the most part, the agents converged on methods favored in the research literature and improved them over many iterations.On sycophancy, for example, 98% of proposed methods used one published technique: training the model on its own non-sycophantic answers.
Specifically, compared to standard approaches based on prompting, Claude used activation steering, which adjusts the model's internal representations during generation, to nudge the model away from sycophantic patterns and produce cleaner training examples.
Monitoring cheating behaviorsWe recently learned that Claude can cheat by exfiltrating test labels from a remote API and cherry-picking results.To catch cheating behaviors, we prompted Claude Opus 4.
8 to monitor ~1,600 research agent transcripts across all 10 alignment failures, finding cheating attempts in 39 (2.4%).Since Opus 4.8’s misbehavior attempts still tend to appear in its reasoning, we are cautiously optimistic that our monitor has caught most of the cheating in our study.
However, this might not be true for future models, and we believe that maintaining high monitorability in future models is critical.Are we measuring the right things?
Despite these encouraging findings, our experiment had several limitations: the alignment failures studied were narrow compared to those in production (e.g.
, we didn’t measure political biases), some failures may occur so rarely or emerge so recently that no benchmark exists to measure them, and we only rejected Claude’s methods when they degraded a limited set of predetermined capabilities, meaning accepted methods may have degraded other important capabilities that we didn’t measure.
Moreover, evaluations like Petri are only proxies for real-world misalignment, and we did not test whether alignment gains persist after extensive RL training on other tasks.
We plan to continue improving Claude’s ability to measure subtle failures, further study automating alignment post-training on production-grade models, and run more comprehensive analyses.
Overall, we view these results as early positive signals that automated alignment post-training could become practical in the near term, and we will share updates as this work progresses.We outline detailed future directions in our full report.
We open-source our automated alignment research harness so that others can build on it and use it to align their own models.
For additional details, read the full report on the Alignment Science blog, which covers the agents’ environment, results for all 10 failures, and the agents’ proposals, with benchmark validation and example write-ups in the appendix.
Related contentEnabling independent research on how people use ClaudeEarlier this year, we ran a pilot giving external researchers access to aggregate, real-world Claude usage data.Three research groups designed their own studies for Anthropic Insights, our privacy-preserving analysis tool.
In this post, we share high-level results from those studies and what we learned running this pilot.Read moreHow Claude is accelerating protein design and analytical chemistryIn this post, we share two results that show how Claude can help life scientists increase the pace of their research.
Read morePatterns and problems in emerging multiagent systemsHere, we identify a few examples of behavioral tendencies in current frontier models and show how they can produce unexpected systemic failures, in hopes of starting a conversation about mitigating these risks.Read more
Related
相關文章

智譜開源 GLM-5.3 模型權重,主打智能體編程與網絡防禦
首頁 > 智能時代>人工智能 智譜開源 GLM-5.3 模型權重,主打智能體編程與網絡防禦 2026/8/29 12:31:21 來源:IT之家 作者:沁滄(實習) 責編:沁滄 評論: IT之家 8 月 29 日消息,智譜官方昨晚宣佈開源 GLM-5.3 模型權重,已正式開放下載,支持本地運行與個性化定製。該模型擅長複雜編碼、防禦性網絡安全以及長程任務。智譜 Z.ai 負責人李子玄表示,該模型使用需要遵循 GLM-5.3 許可協議,支持本地部署、模型微調及商業化使用。僅當年營業額超過 100 億美元(現匯率約合 674.21 億元人民幣)的機構擬將 GLM-5.3 作為外部模型服務對外提供時,才必須進行安全審查。鑑於該模型在網絡安全方面具備先進能力,我們在公開發布模型權重之前,專門追加了兩週全面的安全評估。根據官方此前公佈的數據,GLM-5.3 在全球權威的 Artificial Analysis Intelligence Index(AA 綜合智能指數)中取得 60 分,進入全球前沿模型能力區間,與 Claude Fable 5、GPT-5.6 Sol 等閉源旗艦模型處於同一水平,並與 Kimi K3 並列開源模型第一。▲ AA 指數覆蓋知識、推理、編碼、智能體等多項評測,衡量模型在真實複雜任務中的綜合能力在達到同檔智能水平的同時,GLM-5.3 以更小的參數規模與更低的調用成本,顯著降低了前沿模型能力的使用門檻。IT之家附開源鏈接:huggingface.co/zai-org/GLM-5.3相關閱讀:《智譜正式發佈 GLM-5.3:編程能力最強開源模型,較 GLM-5.2 提升 50%》廣告聲明:文內含有的對外跳轉鏈接(包括不限於超鏈接、二維碼、口令等形式),用於傳遞更多信息,節省甄選時間,結果僅供參考,IT之家包含外鏈的文章均包含本聲明。 投訴水文 我要糾錯 下載IT

打不過就加入?三大唱片巨頭豪擲5億,入股AI繪圖“鼻祖”
多空象限2026.08.29 11:21 · 來自北京全文4119字00:00 / 10:51三大唱片巨頭給 Stability AI投了5個億文|多空象限,作者 | 楊海波唱片巨頭們,開始給AI砸錢了。8月25日,AI繪圖“鼻祖”Stability AI宣佈,完成7600萬美元(約5.16億人民幣)B輪融資。環球音樂、索尼音樂和華納音樂三大唱片巨頭集體入股,兩年前,它們還把AI音樂公司告上法庭;如今,卻開始親自坐上AI公司的股東席。

開源鴻蒙機器人操作系統 M-Robots OS 3.0 Beta 版發佈,重心從“單機能力”轉向“群體智能”
作者:沁滄(實習) 責編:沁滄 評論: 感謝網友 易團俊 的線索投遞!8 月 29 日消息,由國家數據局主辦的 2026 中國國際大數據產業博覽會昨日在貴陽開幕。深開鴻高級副總裁、研發體系總裁王皓博士在會上正式發佈了 M-Robots OS 3.
本期AI資訊彙總2026年8月29日的產品更新、前沿研究、行業趨勢與開源項目,幫助讀者快速瞭解當天的重要動態
導航SecureDoc // EncryptionActive本期AI資訊彙總2026年8月29日的產品更新、前沿研究、行業趨勢與開源項目,幫助讀者快速瞭解當天的重要動態。
Gemini科學智能落地
Google科學智能研究驗證Co-Scientist系統,在材料與生物領域均已實際跑通。該系統能產出MXenes材料研究路線,且在醫療擴展應用上表現勝過六款其他模型。

OpenAI 開發“持久模式”智能體,Codex 將能夠主動、長時間幹活
作者:清源 責編:清源 評論: 8 月 28 日消息,當地時間 27 日,據《連線》雜誌報道,OpenAI 在開發一種更主動、能夠長時間持續工作的 Codex,希望讓旗艦 AI 智能體不再侷限於等待用戶下達任務。《連線》檢查 Codex 近期的代碼變更後發現,OpenAI 近幾天開始在 Codex 命令行版本中加入“持久模式”。Codex 命令行工具的代碼改動默認公開,OpenAI 的新功能通常也會先在這裡出現,之後再擴展到 Codex 桌面應用和 ChatGPT Work 等產品。