Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
Humanity's Last Exam (HLE) is a benchmark dataset of 2,500 expert-crafted questions designed to evaluate the capabilities of advanced AI systems. It was developed through a collaboration between the Center for AI Safety and Scale AI, with contributions from hundreds of researchers worldwide.
Parse Score
Sources
getmaxim.ai shapes more of what AI says about Humanity's Last Exam than any other source, at 33% of its citations.
iternal.ai · medium.com