Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
MedHallu is a benchmark dataset of 10,000 medical question-answer pairs designed to detect hallucinations in large language model outputs. It categorizes hallucinations into easy, medium, and hard tiers, and tests models like GPT-4o and Llama 3.1 on their ability to identify factually incorrect medical responses.
Parse Score