Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
LiveBench is a contamination-free benchmark and leaderboard for evaluating large language models (LLMs) using questions that have verifiable, objective ground-truth answers. It currently contains 21 diverse tasks across 7 categories and refreshes its questions every six months to minimize test-set contamination, with the latest version being LiveBench-2025-11-25. Sponsored by Abacus.AI, it enables objective model comparison across tasks such as reasoning, coding, mathematics, language, and data analysis on an openly accessible leaderboard.
Parse Score