Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
CRAG, short for Comprehensive RAG Benchmark, is a factual question-answering benchmark designed to advance research in retrieval-augmented generation (RAG). It provides mock APIs to simulate web and knowledge-graph search for end-to-end evaluation. The dataset spans five domains and eight question categories with varying entity popularity and temporal dynamics, and includes evaluation metrics that score responses as perfect, acceptable, or miss.
Parse Score