Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
Arena Hard Auto is an automatic evaluation tool for instruction-tuned LLMs that uses automatic judges like GPT-4.1 and Gemini 2.5 to approximate human preferences. It provides a leaderboard ranking models based on performance on hard prompts and creative writing queries sourced from Chatbot Arena.
Parse Score