Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
Hugging Face's Open LLM Leaderboard is a gated collection of datasets that provides evaluation details for large language models. It tracks and compares model performance through community-contributed benchmarks and updated results.
Parse Score