Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
RGB is an implementation for benchmarking large language models in retrieval augmented generation. It provides datasets and evaluation scripts to test information integration, counterfactual robustness, and rejection capabilities.
Parse Score