Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
FACTS Grounding is a benchmark that evaluates the factuality and grounding of large language models by requiring long-form responses grounded in provided source material. It provides a dataset of 1,719 examples across domains such as finance, law, and medicine, split into public and private sets, and uses three evaluators (Gemini 1.5 Pro, GPT-4o, and Claude 3.5 Sonnet) to assess eligibility and factual accuracy with scores averaged. A Kaggle-hosted FACTS leaderboard tracks model progress and the benchmark aims to reduce hallucinations and improve real-world applicability, with ongoing iterations to stay aligned with industry progress.
Parse Score