Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
OckBench is a benchmark that jointly measures accuracy and token efficiency for large language models, introducing a unified metric called OckScore. It reveals that many models consume vastly more tokens than necessary to reach correct answers, with open-source models using up to 5.1× more tokens than commercial counterparts for comparable accuracy.
Parse Score