Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
BenchmarkAggregator is a framework that provides rigorous, unbiased, and scalable evaluations of Large Language Models across diverse AI benchmarks like GPQA Diamond and Chatbot Arena. It offers a unified leaderboard with balanced, cost-efficient comparisons by randomly sampling benchmark questions and integrating models via OpenRouter.
Parse Score