Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
VisualAgentBench (VAB) is a benchmark for evaluating and developing large multimodal models as visual foundation agents across embodied, GUI, and visual design tasks. It provides training trajectories and supports open LLMs/LMMs, building on the AgentBench framework with a focus on visual inputs.
Parse Score
Sources
github.com shapes more of what AI says about VisualAgentBench than any other source, at 100% of its citations.