AAR露出自改進潛力
Training AI models with other AI models has become a very popular goal for neolabs — and now, a researcher in Anthropic’s fellows program has given us an early look at what it might look like in practice.
On Friday, Anthropic published a new paper titled “Automated Researchers Can Reliably Mitigate Alignment Failures,” detailing how AI systems could reliably improve a model’s performance on a set of alignment benchmarks.
When given 10 benchmarks for specific misaligned behaviors, the automated systems were able to improve performance on every single one without degrading overall performance.Led by Anthropic fellow Chen Yueh-Han, the system replicates much of the traditional approach to research.
Each automated system searches the available literature, proposes a method, and trains the model using that method for 30 minutes, gradually increasing the benchmark over several iterations.
Effective methods are preserved while ineffective ones are discarded, allowing the system to operate quickly and at a great scale.“Overall, these results provide early evidence that automated alignment post-training could become practical in the near term,” the paper reads.
The paper is a step toward recursive self-improvement, which many see as the next significant step in AI progress.If models can improve their own alignment training, it’s plausible they could improve training practices more broadly — at which point, human AI researchers might soon become obsolete.
The paper isn’t shy about addressing this idea, explicitly comparing the Automated Alignment Researcher (AAR) to its human equivalent.“The best AAR method beats what experienced humans propose, on average within six hours,” the paper reads.
“Human guided research directions do not lead to stronger performance.” There’s even a cost comparison, in case anyone wasn’t convinced.“An AAR costs roughly $4 per hour in API inference against the $150 per hour we pay our human researchers.
” In fairness, the paper also points out a few limitations to this approach.
The automated system only works insofar as the benchmarks reflect the actual alignment goals, and even then there’s significant work to be done in establishing and maintaining those benchmarks — not to mention maintaining and expanding on the literature the automated researchers are drawn from.
Topics AI, Anthropic When you purchase through links in our articles, we may earn a small commission.This doesn’t affect our editorial independence.Russell Brandom AI Editor Russell Brandom has been covering the tech industry since 2012, with a focus on platform policy and emerging technologies.
He previously worked at The Verge and Rest of World, and has written for Wired, The Awl and MIT’s Technology Review.He can be reached at [email protected] or on Signal at 412-401-5489.View Bio October 13 – 15 San Francisco Don’t miss out.
The startup community will gather to answer a pivotal question: How do you build sustainably in the AI era?
REGISTER NOW Most Popular Nvidia closes in on Hugging Face acquisition Connie Loizos Fitbit founders launch Luffu Link, an LTE health and safety band Aisha Malik Hugging Face reportedly in talks to be acquired for $13B Rebecca Bellan Who’s behind the new ‘stealth model’ Ox Alpha?
Anthony Ha Flock CEO calls for ‘compromise’ as surveillance company faces growing backlash Anthony Ha Two years after launch, Walmart’s Flipkart is closing in on India’s quick-commerce leaders Jagmeet Singh Inherent, founded by DeepMind alumni, says its AI ‘teammate’ just outperformed Anthropic and OpenAI at replicating research Anna Heim
Related
相關文章
Keras解碼宇宙射線
Keras官方科學部落格近日發表最新應用,展示深度學習在宇宙射線研究上的突破。傳統上,分析宇宙射線需要從複雜的觀測資料中手動提取特徵,但這次研究團隊選擇讓模型直接讀取原始波形,跳過繁瑣的前處理環節。透過陣列探測器,系統能進一步提取出由13×13個站點組成的訊號圖,完整保留空間分佈資訊,作為神經網路的輸入。 這套以Keras打造的模型,核心任務是同時判斷宇宙射線抵達時的能量與來向。不同於過去仰賴人工判讀或簡化統計方法,端到端的深度學習架構能自動捕捉波形中細微的變化模式,讓事件重建的精確度獲得提升。

靈掌機器人完成數千萬元天使輪融資
首頁 IT圈 最會買 設置 日夜間 隨系統 淺色 深色 主題色 黑色 投稿 訂閱 RSS訂閱 收藏 軟媒應用 App客戶端 要知App 軟媒魔方 業界 手機 電腦 測評 視頻 AI 蘋果 iPhone 鴻蒙 軟件 智車 數碼 學院 遊戲 直播 5G 微軟 Win10 Win11 專題 搜索 首頁 > 智能時代>具身智能 靈掌機器人完成數千萬元天使輪融資 2026/8/28 21:09:19 作者:潞源 責編:潞源 評論: 8 月 28 日消息,據《科創板日報》今天報道,工業具身機器人技術企業靈掌機器人近日完成數千萬元天使輪融資,投資方包括凱龍高科、遨博智能、錫創投、芯能創投。
GLM-5.3-Flash vs Qwen3.8-Flash-Next: Two Chinese AI Labs Independently Converge on the Same Model Architecture
Two frontier open-weight models shipped within a day of each other this week.Z.ai released GLM-5.3-Flash, a 320B-parameter multimodal MoE model with 18B active parameters.

騰訊 WorkBuddy,學不會豆包工作?
藍洞商業2026.08.28 16:25 · 來自雲南全文4561字00:00 / 10:45騰訊 WorkBuddy 會整合企業微信嗎?文 | 藍洞商業,作者 | 趙衛衛字節跳動的豆包工作上線當天,騰訊雲與智慧產業事業群 CEO 湯道生髮布了一篇內部信長文。Chatbot 的戰役已經是過去時了, AI 辦公智能體才是當下的熱點。湯道生提到了過去一年元寶與豆包的競爭,承認對手值得學習,而元寶在巨大的用戶增長壓力下,花了較多精力去做推廣與引流,當時模型與產品都還沒 ready,其實效果並不滿意。

售價2681元,抱抱臉開源機器人來了
機器人前瞻(公眾號:robot_pro) 作者 | 周加琦 編輯 | 漠影 機器人前瞻8月28日報道,昨天,Hugging Face發佈了一款鴨形開源機器人Microduck,售價399美元(約2681.28元),目前預購已開啟。 Hugging Face聯合創始人兼CEO Clem Delangue在X上稱,歡迎來到開源且平價的機器人時代,物理智能與世界模型技術將實現全民普惠! ▲(圖源:X) 這是Hugging Face推出的第二款機器人。
