Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
This repository provides code for the paper "Judging the Judges," which evaluates alignment and vulnerabilities in LLMs acting as judges. It uses TriviaQA as a benchmark to assess objective knowledge reasoning across nine judge models and nine exam taker models.
Parse Score