Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
AgentVista is a benchmark for evaluating generalist multimodal agents across 25 sub-domains that require long-horizon tool use grounded in realistic visual scenarios. The benchmark reveals that even the best model achieves only 27.3% overall accuracy, demonstrating significant gaps in current agents' ability to perform multi-step reasoning with hybrid tools.
Parse Score