Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
MIRAGE Bench is a multilingual benchmark for evaluating Retrieval Augmented Generation (RAG) systems across 18 languages. It combines heuristic features with a surrogate judge model to produce cost-effective and accurate leaderboards that correlate highly with expensive LLM-based evaluations.
Parse Score
Sources
arxiv.org shapes more of what AI says about MIRAGE-Bench than any other source, at 100% of its citations.