Constitutional AI: Harmlessness from AI Feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell et al. · Anthropic (Technical Report) · 2022
Abstract
Replaces most of the human labelling in harmlessness training with a written set of principles. The model critiques and revises its own responses against that constitution, and those revisions become the training data; a second stage has the model rank response pairs itself, producing a preference signal without human raters in the loop.
Why it matters
Notable for making the values explicit and auditable rather than implicit in whichever labels the annotators happened to produce. The self-critique loop is also the ancestor of the LLM-as-judge pipelines now used far beyond safety work.
https://arxiv.org/abs/2212.08073