Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
WeVal is an open platform for building transparent, reproducible qualitative benchmarks to evaluate AI models, developed by a community of over 1,000 contributors. It provides leaderboards measuring safety, fluency, and helpfulness by averaging results across multiple evaluations, such as those for hallucination, bias, and cultural competency.
Parse Score
Sources
aisecurityandsafety.org shapes more of what AI says about Weval than any other source, at 50% of its citations.
civiceval.org