Library

Research Library

Preprint2026

Pass the Baton: Trajectory-Relayed On-Policy Distillation

Haolei Xu, Xiaowen Xu, Haiwen Hong, Zixuan Ni et al. · arXiv · 2026

Abstract

On-policy distillation trains a student on its own output, which breaks when the student commits early to a wrong direction: everything after that builds on the mistake, and teacher feedback on it is worthless. The authors detect those moments by noticing that teachers redirect where students plough on, hand control to the teacher briefly, then give it back — concentrating a limited intervention budget on early, high-leverage positions.

Why it matters

Training small models from larger ones is now a standard production step. This is a cheap, label-free fix for the failure mode that wastes most of that compute.

distillationtrainingreasoningsmall models
Read the source

https://arxiv.org/abs/2607.26057