Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
RobustJudge is a fully automated, scalable evaluation framework that systematically tests the robustness of LLM-as-a-Judge systems against adversarial attacks and defense strategies. It runs diverse attack types, applies and measures defenses, and reports metrics like SDR/iSDR and ASR to quantify how easily judges can be manipulated.
Parse Score
Sources
arxiv.org shapes more of what AI says about RobustJudge than any other source, at 100% of its citations.