Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
ResearchCodeBench is a benchmark that evaluates large language models on their ability to generate executable code from novel machine learning research papers published in 2024–2025. It comprises 212 coding challenges from top ML papers and finds that even the best models solve fewer than 40% of the tasks, with most challenges unseen during pretraining.
Parse Score