Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
Safety J is an advanced bilingual (English and Chinese) safety evaluator that assesses content generated by Large Language Models (LLMs) through detailed critiques and safety classifications. It outperforms existing open-source and proprietary models like GPT-4o in safety evaluation tasks, using an iterative preference learning process and a meta-evaluation framework.
Parse Score