Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
CompassJudger 2 is a generalist judge model that evaluates large language models using verifiable rewards and rejection sampling to enhance judgment accuracy and robustness. It achieves competitive performance across multiple benchmarks, with its 7B variant matching larger models like DeepSeek V3 and Qwen3 235B.
Parse Score