DRIFT大模型自進化框架發佈
Computer Science > Machine Learning arXiv:2606.
30345 (cs) [Submitted on 29 Jun 2026 (v1), last revised 10 Aug 2026 (this version, v4)] Title:DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training Authors:Haisen Luo, Yiwei Liu, Haoning Wang, Dan Liu, Junxi Yin, Haotian Wang, Lei Zhang, Xiaoyu Tian, Shuaiting Chen, Yuansheng Song, Baoyan Guo, Xiongfei Yan, Bolan Yang, Chengwei Liu, Ming Cui, Jiong Chen View a PDF of the paper titled DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training, by Haisen Luo and 15 other authors View PDF HTML (experimental) Abstract:Enabling large language models to achieve stable self-improvement without external expert supervision remains a central challenge in complex reasoning tasks.
Existing self-distillation and reinforcement learning methods lack explicit mechanisms for tracking problem-level learning progress and adapting optimization strategies accordingly.
Consequently, training may over-optimize easy problems, receive weak supervision from hard problems, and fail to sufficiently explore borderline cases.To resolve these issues, we propose DRIFT, an online self-evolution policy optimization framework for large language models.
DRIFT regulates the model's self-improvement process through the joint use of Difficulty Routing and Rhythm Gating.
The former identifies the model's learning state at the problem level and dynamically allocates self-distillation and reinforcement learning signals, while the latter refines policy updates at the token level, concentrating exploration on critical reasoning positions.
By further incorporating a success buffer and a two-stage curriculum learning strategy, DRIFT preserves high-quality historical experience while progressively guiding the model from reliable behavior acquisition toward stable policy evolution.
Evaluated across five benchmarks and three model scales, DRIFT surpasses the peak performance of both GRPO and SDPO across all evaluated metrics.On the average score over the five benchmarks, DRIFT achieves 79.5$\%$, outperforming GRPO by 9.5$\%$ and SDPO by 7.
5$\%$, establishing a new state-of-the-art result.Notably, on ToolUse, DRIFT reaches an accuracy of 79.2$\%$, improving over GRPO by 13.5$\%$ and SDPO by 10.7$\%$, setting a new state-of-the-art and substantially outperforming all concurrent methods.Subjects: Machine Learning (cs.
LG); Artificial Intelligence (cs.AI) Cite as: arXiv:2606.30345 [cs.LG] (or arXiv:2606.30345v4 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2606.
30345 Focus to learn more arXiv-issued DOI via DataCite Submission history From: Haoning Wang [view email] [v1] Mon, 29 Jun 2026 14:20:47 UTC (5,472 KB) [v2] Fri, 31 Jul 2026 04:05:32 UTC (5,574 KB) [v3] Mon, 3 Aug 2026 04:06:03 UTC (5,574 KB) [v4] Mon, 10 Aug 2026 05:19:37 UTC (5,574 KB) Full-text links: Access Paper: View a PDF of the paper titled DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training, by Haisen Luo and 15 other authorsView PDFHTML (experimental)TeX Source view license Current browse context: cs.
LG < prev | next > new | recent | 2026-06 Change to browse by: cs cs.AI References & Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading...BibTeX formatted citation × loading...
Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.
ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?
) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.
AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?) IArxiv recommender toggle IArxiv Recommender (What is IArxiv?
) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy.arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community?Learn more about arXivLabs.Which authors of this paper are endorsers?| Disable MathJax (What is MathJax?)
Related
相關文章

偷書要賠15億美元,AI巨頭燒幾百萬本書反而合法了
偷書要賠15億美元,AI巨頭燒幾百萬本書反而合法了藍字計劃2026.08.24 18:01 · 來自廣東全文4120字00:00 / 11:49只要能訓練出更好的大模型,究竟要銷燬多少本實體書,甚至在實體書之外還會消耗多少前AI時代的“資產”,也許就再也沒人關心了。文 | 藍字計劃,作者|Chester一場現代版的“焚書”,正在美國發生。最近幾年,美國二手書市場出現了一個奇怪現象:不少神秘買家,一次性採購成百上千本書,不挑書、不問價,甚至連冷門舊書都照買不誤。隨著電子閱讀的普及,現在還有多少人看實體書?更別說這種成百上千的掃貨始買書。書商們自然也開始好奇:究竟是誰在買走這批書?最近,美國科技媒體404 Media聯繫到一名二手書商。對方剛剛通過書籍交易平臺Biblio,接到了一筆大約1000本舊書的訂單,其中還包括不少稀有、絕版和具有收藏價值的書。為了找出背後的買家,404 Media把一枚AirTag塞進其中一本書裡,然後一路追蹤。最終,這本書被送進了亞馬遜位於拉斯維加斯的一家倉庫。據404 Media調查,這些書到了亞馬遜倉庫後,命運卻只有三步:切掉書脊、批量掃描、然後扔進碎紙機。負責這項工作的VGT3團隊,甚至還給自己設計了一個頗為應景的Logo:一隻露著牙齒、手裡抓著書的霸王龍。 自己也賣書的亞馬遜卻幹起了銷燬實體書的活已經夠驚奇了,但更驚奇的是同樣這樣乾的,不只有亞馬遜。在更早之前,Anthropic就被曝光過一個代號為“巴拿馬計劃”的秘密項目。他們的目標,是把全世界的書都“破壞性掃描”一遍。在大約一年時間裡,Anthropic為此花費數千萬美元,買下數百萬本實體書,然後切掉書脊、掃描內容,再把剩下的紙張送去回收。美國的AI大公司,怎麼就和舊書幹上了?互聯網內容,不夠AI用了AI公司瘋狂買二手書,其實是想買下書裡的內容和數據。過去,互聯網是大模型最方便的數據來源。

阿里視頻大模型Wan3.0正式上線,行業評價“穩定、真實、有質感”
阿里巴巴影片生成大模型Wan3.0正式上線,單次可生成30秒影片,並首次支援doc、xls、ppt、pdf、md等文檔輸入。企業用戶普遍評價其「穩定、真實、有質感」,能穩定保持角色與場景一致性,並已進入短劇、影視、廣告等生產流程。即日起可於阿里雲百鍊、千問等平台體驗,標準版並推出限時7折優惠。

月之暗面第一代萬億參數多模態模型 Kimi K2.5 官宣月底結束服役
作者:歸瀧 責編:歸瀧 評論: 8 月 24 日消息,月之暗面 Kimi 官方微博今日宣佈,其第一代萬億參數多模態模型 —— Kimi K2.5 本月底即將結束服役。據此前報道,今年 1 月,月之暗面宣佈推出並開源了其最新的 Kimi K2.

消息稱知名 AI 研究員 Luke Metz 離開 OpenAI,加入 Meta 超級智能實驗室
作者:遠洋 責編:遠洋 評論: 感謝網友 華南吳彥祖 的線索投遞!8 月 24 日消息,據知情人士向 Axios 證實,知名 AI 研究員 Luke Metz 已加入 Meta 的超級智能實驗室(Superintelligence Labs)。

Anthropic 最強大模型 Fable 5 遇冷,企業用戶轉向更便宜 AI 產品
作者:遠洋 責編:遠洋 評論: 8 月 24 日消息,據英國《金融時報》報道,Anthropic 的美國客戶正在使用更便宜的替代品來替代其最強大的 AI 工具,這在其預計將實現有史以來規模最大的 IPO 之前,對其高支出的商業模式提出了質疑。

阿里雲視頻生成模型 Wan3.0 正式上線,支持單次生成 30 秒視頻、文檔輸入
作者:遠洋 責編:遠洋 評論: 8 月 24 日消息,阿里雲消息,今天,視頻生成模型 Wan3.0 正式上線。官方稱,Wan3.0 在生成時長、萬能創作、全能參考以及真實世界還原等維度全面升級,單次可生成 30 秒視頻,並首次支持 doc、xls、ppt、pdf、md 等文檔格式輸入,力求準確還原真實世界。