Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
AlpacaEval is an automatic evaluator for instruction-following language models that uses a powerful LLM to compare model outputs against a reference. It is designed to be fast (under 5 minutes), cheap (under $10), and highly correlated with human judgments (0.98 Spearman correlation with ChatBot Arena).
Parse Score