Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills
Siyuan Huang, Pengyu Cheng, Haotian Liu, Tao Chen et al. · arXiv · 2026
Abstract
Self-improving training loops face a trade-off: environments give reliable feedback but cover narrow domains, while freely invented tasks cover everything and verify nothing. This work proposes agent skills as the middle ground — each skill is narrow enough to verify execution against, while routing across a growing library of them keeps the task distribution open-ended. A proposer, a solver, and a skill controller co-evolve inside one reinforcement-learning loop.
Why it matters
Self-play is how systems surpassed human data in games, and this is a serious attempt to port that to general model capability. Read it alongside the other co-evolution work here to see the pattern forming.
https://arxiv.org/abs/2607.22529