Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
OSWorld is a scalable, real computer environment for benchmarking multimodal agents on open-ended tasks across operating systems like Ubuntu, Windows, and macOS. It provides 369 real-world computer tasks with reproducible setup and evaluation scripts, revealing that current AI agents achieve only 12.24% success compared to 72.36% for humans.
Parse Score
Sources
arxiv.org shapes more of what AI says about OSWorld than any other source, at 100% of its citations.