Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
OmniBench is a self-generating, graph-based benchmark for evaluating multimodal large language model (MLLM) virtual agents across 10 capabilities. It uses an automated pipeline to synthesize 36,000 graph-structured tasks with controllable complexity across 20 scenarios, achieving a 91% human acceptance rate.
Parse Score