TL;DR
An empirical study of how optimizing against a learned reward model can eventually reduce performance under a higher-quality reference signal.
Why it matters
Reward hacking is a predictable scaling risk. Stronger optimization is not automatically better when the proxy reward is imperfect.
Key findings
- 01
Proxy-reward gains can diverge from the reference objective under continued optimization.
- 02
Reward-model quality and optimization strength must be treated as coupled scaling variables.
Scaling dimensions
Models, methods & benchmarks
- Algorithms
- Reward model optimization
Topics