Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
Quesma provides independent audits of agentic scaffolding and harnesses for AI coding benchmarks, comparing how agent workflows and test verification impact performance. Their analysis evaluates tools like Blitzy against base models such as GPT 5.4, Gemini 3.1 Pro, and Claude Code on SWE Bench Pro.
Parse Score