Library

Research Library

Preprint2026

DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space

Jiangwang Chen, Zixin Song, Junlin Liu, Shuaiyu Zhou et al. · arXiv · 2026

Abstract

A model optimised against a fixed rubric eventually saturates the criteria that rubric measures and stops improving on everything it ignores. Evolving the rubric alongside the solver is the obvious fix, except that scoring rubric updates by the solver's own score quietly rewards making the rubric easier. This work separates the two objectives: the solver learns from criterion-level feedback while the rubric generator is audited for coverage and discrimination independently.

Why it matters

Relevant well past research: anyone running an LLM-as-judge pipeline is optimising against a fixed rubric and will meet the same ceiling. The decoupling argument is the part to take away.

evalsllm as judgereinforcement learningoptimization
Read the source

https://arxiv.org/abs/2607.25675