Data as of Sep 26, 2026 · Based on 4,029,442 AI responses across 13,338 prompts · See how Parse measures this
EleutherAI's lm-evaluation-harness is an open-source framework for evaluating language models across a wide range of tasks and benchmarks, including Open LLM Leaderboard-style datasets. It offers a command-line interface and YAML-based configuration to organize tasks and run evaluations, with support for multiple backends such as Hugging Face transformers and VLLM for batched and asynchronous requests. The project includes a growing set of task definitions, experimental multimodal capabilities, and tooling to compare model performance on standardized tasks and leaderboards.
Parse Score
No contexts measured yet.