Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
Judge Arena is a platform for benchmarking large language models (LLMs) as evaluators, using community rankings to assess their performance. It leverages Together AI for inference of open-source models, with FP8
Parse Score
Sources
huggingface.co shapes more of what AI says about Judge Arena than any other source, at 100% of its citations.