Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
Vectara's Hallucination Leaderboard is a public benchmark that evaluates how often large language models (LLMs) introduce hallucinations when summarizing documents. It uses Vectara's Hallucination Evaluation Model (HHEM) to compute hallucination rates and factual consistency scores for various LLMs.
Parse Score
Sources
en.wikipedia.org shapes more of what AI says about Hallucination Leaderboard than any other source, at 50% of its citations.
huggingface.co