Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
SWE bench is a benchmark for evaluating large language models on real-world software engineering tasks, offering subsets like Verified, Lite, Multilingual, and Multimodal. It provides leaderboards and metrics such as the percentage of resolved instances to compare model performance.
Parse Score
Sources
aimultiple.com shapes more of what AI says about SWE-bench than any other source, at 9.1% of its citations.
anthropic.com · arxiv.org · demandsphere.com · iternal.ai