DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space
Jiangwang Chen, Zixin Song, Junlin Liu, Shuaiyu Zhou et al. · arXiv · 2026
Abstract
A model optimised against a fixed rubric eventually saturates the criteria that rubric measures and stops improving on everything it ignores. Evolving the rubric alongside the solver is the obvious fix, except that scoring rubric updates by the solver's own score quietly rewards making the rubric easier. This work separates the two objectives: the solver learns from criterion-level feedback while the rubric generator is audited for coverage and discrimination independently.
Why it matters
Relevant well past research: anyone running an LLM-as-judge pipeline is optimising against a fixed rubric and will meet the same ceiling. The decoupling argument is the part to take away.
https://arxiv.org/abs/2607.25675