The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation
Tianyi Men, Zhuoran Jin, Kang Liu, Jun Zhao · arXiv · 2026
Abstract
A controlled study of where an agent's planning ability actually comes from, run inside a purpose-built environment so the training data can be varied deliberately instead of inherited from the open web. It separates three stages: which pretraining data produces long-horizon generalisation, how post-training reshapes it, and how much of the result is a general planning pattern versus task-specific knowledge. Mediocre training trajectories turn out to be actively harmful, because errors compound over a long horizon.
Why it matters
Most agent papers report that a technique works; this one tries to explain why. The finding that suboptimal trajectories hurt more than they help is a direct warning about training on scraped agent traces.
https://arxiv.org/abs/2607.24720