Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
HallusionBench is an advanced diagnostic suite for evaluating entangled language hallucination and visual illusions in large vision-language models. It provides an image-context reasoning benchmark designed to challenge multimodal models (e.g., GPT-4V, LLaVA-1.5) and expose failures in how models align visual input with language outputs. The repository includes datasets, evaluation scripts, and a leaderboard to reproduce experiments and compare model performance.
Parse Score