TL;DR
This work studies language models that generate candidate responses and also provide the preference signal used to improve subsequent iterations.
Why it matters
If model-generated feedback is reliable enough, the feedback loop can scale with less direct human labeling—but it also concentrates evaluation risk in the model itself.
Key findings
- 01
Iterative self-improvement depends on both response generation and judge quality.
- 02
Self-reward creates a scalable feedback path while raising calibration and bias questions.
Scaling dimensions
Models, methods & benchmarks
- Algorithms
- Iterative preference optimization
Topics