Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
ZeroSumEval is an extensible framework for evaluating large language models (LLMs) through competitive, multi-agent games that scale in difficulty as models improve. It pits models against each other in structured games with clear win conditions to assess knowledge, reasoning, planning, and other capabilities, using DSPy-based optimization to evaluate self-improvement while keeping competition fair. The project provides a growing evaluation suite and configuration for running these dynamic benchmarks rather than fixed, subjective assessments.
Parse Score