Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
BEIR is a heterogeneous benchmark for zero-shot evaluation of information retrieval models, providing a common framework to evaluate lexical, dense, sparse, and reranking architectures. It includes 17 preprocessed datasets and supports easy integration of custom models with state-of-the-art evaluation metrics.
Sources
arxiv.org shapes more of what AI says about BEIR than any other source, at 17% of its citations.
benchmarkingagents.com · datarobot.com · getmaxim.ai · python.plainenglish.io
The market map
Embedding Model APIs and Services →Where AI ranks BEIR