TL;DR

An empirical study of how optimizing against a learned reward model can eventually reduce performance under a higher-quality reference signal.

Why it matters

Reward hacking is a predictable scaling risk. Stronger optimization is not automatically better when the proxy reward is imperfect.

Key findings

  1. 01

    Proxy-reward gains can diverge from the reference objective under continued optimization.

  2. 02

    Reward-model quality and optimization strength must be treated as coupled scaling variables.

Scaling dimensions

Models, methods & benchmarks

Algorithms
Reward model optimization

Topics

Read the original sourcearXiv