Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
LawBench is a comprehensive evaluation benchmark designed to assess the legal knowledge and task performance of large language models (LLMs) within the Chinese legal system. It includes 20 tasks across three cognitive levels—memory, understanding, and application—and reports results for 51 LLMs, with a unique 'abstention rate' metric to measure model refusal or misunderstanding.
Parse Score