Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
CheckList is a tool for behavioral testing of Natural Language Processing (NLP) models, as described in the ACL 2020 paper 'Beyond Accuracy'. It provides a systematic methodology to evaluate NLP models beyond standard accuracy metrics by testing specific linguistic capabilities.
Parse Score