Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
KILT is a benchmark for knowledge-intensive language tasks, providing a standardized evaluation framework for models that require external knowledge. It includes a knowledge source based on Wikipedia and datasets for tasks like fact checking, entity linking, and slot filling.
Parse Score