Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
DoVer is an intervention-driven debugging framework for LLM-based multi-agent systems that augments hypothesis generation with active verification through targeted interventions, such as editing messages or altering plans, to address debugging failures beyond log-based analysis. Rather than focusing on attribution accuracy, it evaluates progress toward task success, measuring whether failures are resolved or milestones are achieved. On datasets derived from GAIA and AssistantBench within the Magnetic-One agent framework, DoVer converts 18-28% of failed trials to successes, achieves up to 16% milestone progress, validates or refutes 30-60% of hypotheses, and also performs well on GSMPlus/AG2 with 49% recovery of failed trials.
Parse Score