TL;DR
Eureka uses a code-capable language model to propose and iteratively improve executable reward functions for reinforcement-learning tasks.
Why it matters
Reward engineering is a major scaling bottleneck. Automating part of it can expand the task frontier, provided generated rewards are evaluated against real behavior.
Key findings
- 01
Executable reward code gives language models a concrete optimization surface.
- 02
Simulation feedback can guide iterative reward improvement without updating the language model.
Scaling dimensions
Models, methods & benchmarks
- Algorithms
- Reward design, Reinforcement learning
Topics