Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
FaithJudge is a benchmark that evaluates hallucinations generated by large language models in Retrieval Augmented Generation tasks, focusing on summarization, question answering, and data-to-text generation. It provides a leaderboard ranking models by their overall hallucination rate across these tasks.
Parse Score