Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
The AI Agent Benchmark Compendium is a curated collection of over 50 modern benchmarks for evaluating AI agents, organized into categories like function calling, reasoning, coding, and computer interactions. It provides links to papers, datasets, and leaderboards for each benchmark, with an open invitation for community contributions via pull requests or issues.
Parse Score