TL;DR

OSWorld introduces a benchmark and environment for evaluating multimodal agents on open-ended tasks across real computer interfaces.

Why it matters

Agent training needs environments whose failures resemble production failures. OSWorld pushes evaluation closer to actual operating systems and applications.

Key findings

  1. 01

    Real computer interaction exposes grounding and long-horizon reliability problems hidden by simpler benchmarks.

  2. 02

    Environment fidelity is central to meaningful agent evaluation.

Scaling dimensions

Models, methods & benchmarks

Benchmarks
OSWorld

Topics

Read the original sourcearXiv