Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
Auto Arena of LLMs is a fully automated evaluation framework that tests large language models through agent peer battles and committee discussions. The system achieves a 94.5% correlation with human preference scores on Chatbot Arena models, outperforming all existing benchmarks.
Parse Score