跳转到内容

On-Policy Self-Distillation for Reasoning Models

作者: Self-Distilled Reasoner Team (2026)

arXiv: 2601.18734

领域

对齐推理能力

TLDR(中文)

2026 开年后训练热点"自蒸馏"的代表工作:模型在自身策略分布上生成数据并蒸馏回自身,不再依赖外部强教师模型,配合可验证奖励实现自我改进闭环,显著降低后训练对蒸馏管线的依赖。

TLDR (English)

A representative work of the early-2026 "self-distillation" trend in post-training: the model generates data on its own policy distribution and distills it back into itself, removing the need for an external strong teacher and closing a self-improvement loop with verifiable rewards, reducing post-training dependence on distillation pipelines.

出现在这些文章里

同被引用

这些论文与本文出现在同一篇文章中

相关论文

同一领域的其他论文