Library

Research Library

Preprint2026

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation

Tianyi Men, Zhuoran Jin, Kang Liu, Jun Zhao · arXiv · 2026

Abstract

A controlled study of where an agent's planning ability actually comes from, run inside a purpose-built environment so the training data can be varied deliberately instead of inherited from the open web. It separates three stages: which pretraining data produces long-horizon generalisation, how post-training reshapes it, and how much of the result is a general planning pattern versus task-specific knowledge. Mediocre training trajectories turn out to be actively harmful, because errors compound over a long horizon.

Why it matters

Most agent papers report that a technique works; this one tries to explain why. The finding that suboptimal trajectories hurt more than they help is a direct warning about training on scraped agent traces.

agentsplanningpretrainingpost training
Read the source

https://arxiv.org/abs/2607.24720