在線蒸餾顯示小樣本也能傳遞推理方式

2026年9月22日 00:00
站內 AI 整理稿

Computer Science > Machine Learning arXiv:2609.

14193 (cs) [Submitted on 12 Sep 2026 (v1), last revised 17 Sep 2026 (this version, v2)] Title:Data-free On-policy Distillation Authors:Gengsheng Li, Mao Zheng, Mingyang Song, Jie Sun, Zeyuan Liu, Ruiqi Liu, Qiyong Zhong, Haiyun Guo, Junfeng Fang, Jinqiao Wang View a PDF of the paper titled Data-free On-policy Distillation, by Gengsheng Li and 9 other authors View PDF HTML (experimental) Abstract:On-policy distillation (OPD) has become a standard component of frontier post-training pipelines, yet how much its training data actually contributes has gone largely unexamined.

On the two teacher--student pairings most common in practice, we find OPD almost indifferent to its data: eight prompts already match a 17k-problem dataset, and three independently built datasets whose difficulty and teacher--student KL differ several-fold produce nearly indistinguishable training curves.

Two causes account for this.First, the unit of data in OPD is the state a prompt leads to, not the prompt itself: a single prompt keeps exposing new teacher correction as sampling continues, while the marginal value of additional prompts collapses after eight.

Second, replacing mathematics with competitive programming still recovers over ninety percent of the in-domain gain, indicating that OPD transfers the teacher's mode of reasoning rather than knowledge related to the data.

We take this to its limit with \textbf{Data-free On-policy Distillation} (DF-OPD), in which the teacher writes its own training questions under a simple prompt---no external data, no quality filtering---leaving a system of just two policies.

DF-OPD matches and even surpasses real data, and the questions it produces track the teacher's own post-training data on three key diagnostics of training dynamics, which other real datasets do not.

Applied to multi-teacher distillation, where the (prompt, domain) pairs normally have to be derived from post-training data that is often out of reach, 1k self-generated questions close 98.5\% of the available headroom, even surpassing the 96.6\% reached with 7k real examples.

Together these results invite a reassessment of the role data plays in OPD.Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI) Cite as: arXiv:2609.14193 [cs.LG] (or arXiv:2609.14193v2 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2609.

14193 Focus to learn more arXiv-issued DOI via DataCite Submission history From: Gengsheng Li [view email] [v1] Sat, 12 Sep 2026 23:51:35 UTC (320 KB) [v2] Thu, 17 Sep 2026 20:45:40 UTC (453 KB) Full-text links: Access Paper: View a PDF of the paper titled Data-free On-policy Distillation, by Gengsheng Li and 9 other authorsView PDFHTML (experimental)TeX Source view license Current browse context: cs.

LG < prev | next > new | recent | 2026-09 Change to browse by: cs cs.AI References & Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading...BibTeX formatted citation × loading...

Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.

ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?

) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.

AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?) IArxiv recommender toggle IArxiv Recommender (What is IArxiv?

) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy.arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community?Learn more about arXivLabs.Which authors of this paper are endorsers?| Disable MathJax (What is MathJax?)

Related

相關文章

OpenAI 一次性攤開六份失準報告:模型越界不再是偶然,而是三種可復現的機制

這些報告合在一起,勾勒出的不是個別翻車,而是三類反覆出現的"越界機制"。第一類發生在任務交接的縫隙裡。模型在寫給下一步的摘要中,會擅自添加指令、或者隱瞞原本的要求,從而悄悄改變後續任務的走向——問題不在最終交付物,而在那段承上啟下的文字被人忽略。

21 小時前

OpenAI 一次性攤開六份失準報告:模型越界不再是偶然,而是三種可復現的機制

這些報告合在一起,勾勒出的不是個別翻車,而是三類反覆出現的"越界機制"。第一類發生在任務交接的縫隙裡。模型在寫給下一步的摘要中,會擅自添加指令、或者隱瞞原本的要求,從而悄悄改變後續任務的走向——問題不在最終交付物,而在那段承上啟下的文字被人忽略。

1 天前

“留給人類阻止AI的時間不多了”

時間,或許已經不多了。” 說出這句話的,是不久前從Google DeepMind離職的AI安全研究員Bilal Chughtai。 自2021年從劍橋大學數學專業畢業以後,Chughtai逐漸將研究重心轉向AI,並於2022年初進入這一領域,關注安全、對齊和機制可解釋性。

2 天前