Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
SafeBench is a comprehensive framework for evaluating the safety of multimodal large language models (MLLMs), featuring an automatic pipeline that generates 2,300 multimodal harmful query pairs across 23 risk scenarios. It employs a jury deliberation evaluation protocol using collaborative LLMs to provide reliable, unbiased safety assessments, and has tested 15 open-source and 6 commercial MLLMs, revealing widespread safety issues.
Parse Score