TL;DR
Constitutional AI uses written principles and model-generated critiques and preferences to reduce direct dependence on human harmlessness labels during alignment training.
Why it matters
AI feedback is one route to scaling supervision, but its reliability depends on the constitution, the judging model, and careful evaluation of blind spots.
Key findings
- 01
Principles can structure model-generated critique and preference data.
- 02
AI feedback changes the supervision bottleneck rather than removing the need for evaluation.
Scaling dimensions
Models, methods & benchmarks
- Algorithms
- RLAIF
Topics