在線蒸餾顯示小樣本也能傳遞推理方式
Computer Science > Machine Learning arXiv:2609.
14193 (cs) [Submitted on 12 Sep 2026 (v1), last revised 17 Sep 2026 (this version, v2)] Title:Data-free On-policy Distillation Authors:Gengsheng Li, Mao Zheng, Mingyang Song, Jie Sun, Zeyuan Liu, Ruiqi Liu, Qiyong Zhong, Haiyun Guo, Junfeng Fang, Jinqiao Wang View a PDF of the paper titled Data-free On-policy Distillation, by Gengsheng Li and 9 other authors View PDF HTML (experimental) Abstract:On-policy distillation (OPD) has become a standard component of frontier post-training pipelines, yet how much its training data actually contributes has gone largely unexamined.
On the two teacher--student pairings most common in practice, we find OPD almost indifferent to its data: eight prompts already match a 17k-problem dataset, and three independently built datasets whose difficulty and teacher--student KL differ several-fold produce nearly indistinguishable training curves.
Two causes account for this.First, the unit of data in OPD is the state a prompt leads to, not the prompt itself: a single prompt keeps exposing new teacher correction as sampling continues, while the marginal value of additional prompts collapses after eight.
Second, replacing mathematics with competitive programming still recovers over ninety percent of the in-domain gain, indicating that OPD transfers the teacher's mode of reasoning rather than knowledge related to the data.
We take this to its limit with \textbf{Data-free On-policy Distillation} (DF-OPD), in which the teacher writes its own training questions under a simple prompt---no external data, no quality filtering---leaving a system of just two policies.
DF-OPD matches and even surpasses real data, and the questions it produces track the teacher's own post-training data on three key diagnostics of training dynamics, which other real datasets do not.
Applied to multi-teacher distillation, where the (prompt, domain) pairs normally have to be derived from post-training data that is often out of reach, 1k self-generated questions close 98.5\% of the available headroom, even surpassing the 96.6\% reached with 7k real examples.
Together these results invite a reassessment of the role data plays in OPD.Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI) Cite as: arXiv:2609.14193 [cs.LG] (or arXiv:2609.14193v2 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2609.
14193 Focus to learn more arXiv-issued DOI via DataCite Submission history From: Gengsheng Li [view email] [v1] Sat, 12 Sep 2026 23:51:35 UTC (320 KB) [v2] Thu, 17 Sep 2026 20:45:40 UTC (453 KB) Full-text links: Access Paper: View a PDF of the paper titled Data-free On-policy Distillation, by Gengsheng Li and 9 other authorsView PDFHTML (experimental)TeX Source view license Current browse context: cs.
LG < prev | next > new | recent | 2026-09 Change to browse by: cs cs.AI References & Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading...BibTeX formatted citation × loading...
Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.
ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?
) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.
AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?) IArxiv recommender toggle IArxiv Recommender (What is IArxiv?
) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy.arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community?Learn more about arXivLabs.Which authors of this paper are endorsers?| Disable MathJax (What is MathJax?)
Related
相關文章

宇樹科技發佈 Dex5-S 靈巧手:22 自由度、真手 1:1 尺寸,售價 3.99 萬元起
USB×1,通信波特率 6Mbps;控制頻率方面,T1s 以太網 / USB 最高 1000Hz,高速 485 為 50Hz。感知反饋涵蓋關節模式、位置、速度、力矩、溫度、電壓電流等,控制指令則包括關節模式、位置、速度、力矩、剛度係數與阻尼係數。

華為汪濤:華為要打造AI算力底座,只做好一顆芯片遠遠不夠
。 昇騰960DT研發進度比原計劃提前三個季度,單芯片算力實現倍增,960DT支持2 PFLOPS FP8、4 PFLOPS FP4,HBM容量最高288GB、帶寬9.6TB/s。 對應的昇騰960超節點做到4096張NPU卡,最高8EFLOPS FP8算力、超過1PB HBM容量。

OpenAI 一次性攤開六份失準報告:模型越界不再是偶然,而是三種可復現的機制
這些報告合在一起,勾勒出的不是個別翻車,而是三類反覆出現的"越界機制"。第一類發生在任務交接的縫隙裡。模型在寫給下一步的摘要中,會擅自添加指令、或者隱瞞原本的要求,從而悄悄改變後續任務的走向——問題不在最終交付物,而在那段承上啟下的文字被人忽略。
OpenAI 一次性攤開六份失準報告:模型越界不再是偶然,而是三種可復現的機制
這些報告合在一起,勾勒出的不是個別翻車,而是三類反覆出現的"越界機制"。第一類發生在任務交接的縫隙裡。模型在寫給下一步的摘要中,會擅自添加指令、或者隱瞞原本的要求,從而悄悄改變後續任務的走向——問題不在最終交付物,而在那段承上啟下的文字被人忽略。

“工作垃圾”飛來飛去,曾大力擁抱 AI 的科技大佬開始“踩剎車”了
作者:清源 責編:清源 評論: 9 月 19 日消息,當地時間 17 日,據《財富》雜誌網站報道,曾經把 AI 吹捧為釋放員工全部潛力關鍵的科技行業領導者,正開始逐漸改變說法。Shopify 聯合創始人兼 CEO 托比亞斯 · 呂特克就是其中之一。
“留給人類阻止AI的時間不多了”
時間,或許已經不多了。” 說出這句話的,是不久前從Google DeepMind離職的AI安全研究員Bilal Chughtai。 自2021年從劍橋大學數學專業畢業以後,Chughtai逐漸將研究重心轉向AI,並於2022年初進入這一領域,關注安全、對齊和機制可解釋性。