Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
WildGuard is an open-source safety classification model for LLM chat exchanges, capable of classifying prompt harmfulness, response harmfulness, and whether a response is a refusal. It supports both prompt-only and prompt-response inputs, and can be used with VLLM or HuggingFace backends.
Parse Score
Sources
arxiv.org shapes more of what AI says about WildGuard than any other source, at 100% of its citations.