Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
AgentBench is a benchmark for evaluating large language models as autonomous agents across diverse environments, including operating systems, databases, knowledge graphs, and web tasks. It provides standardized test splits and leaderboards to measure multi-turn interaction performance of LLMs.
Parse Score