Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
Agent X is a benchmark for evaluating deep multimodal reasoning and tool use in vision-centric LLM agents, featuring 828 authentic tasks across six domains. It assesses step-by-step decision-making, reasoning coherence, and final outcomes, revealing that even top models solve fewer than half of the tasks.
Parse Score