Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
CapaBench is an evaluation framework that uses Shapley Value from cooperative game theory to measure individual module contributions in modular Large Language Models. It assesses Planning, Reasoning, Action, and Reflection capabilities across over 1,500 multi-round task scenarios to enable systematic optimization and interpretability of agent architectures.
Parse Score