Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
JudgeBench is a benchmark for evaluating LLM-based judges, containing 620 unique response pairs generated by GPT-4o and Claude 3.5 Sonnet. It provides source code, datasets, and support for multiple prompted judges and reward models to assess the performance of AI judges.
Parse Score
Sources
huggingface.co shapes more of what AI says about JudgeBench than any other source, at 100% of its citations.