TL;DR

Eureka uses a code-capable language model to propose and iteratively improve executable reward functions for reinforcement-learning tasks.

Why it matters

Reward engineering is a major scaling bottleneck. Automating part of it can expand the task frontier, provided generated rewards are evaluated against real behavior.

Key findings

  1. 01

    Executable reward code gives language models a concrete optimization surface.

  2. 02

    Simulation feedback can guide iterative reward improvement without updating the language model.

Scaling dimensions

Models, methods & benchmarks

Algorithms
Reward design, Reinforcement learning

Topics

Read the original sourcearXiv