Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
LLM AggreFact is a fact-checking benchmark that aggregates 11 publicly available datasets for evaluating grounded factuality and hallucination in large language models. It provides a leaderboard comparing model performance across multiple datasets, such as CNN XSum and RAG Truth.
Parse Score