Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
WorkArena is a benchmark of 33 browser-based tasks built on the ServiceNow platform to measure how well large language model agents can perform common knowledge work tasks. It introduces BrowserGym, an environment for designing and evaluating web agents, and finds that current agents achieve only 42.7% success with GPT-4, leaving significant room for improvement.
Parse Score