Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
AgentRewardBench evaluates automatic evaluators for web agent trajectories, providing a library of environments, agents, and metrics. It includes a leaderboard ranking LLM judges and a dataset hosted on Hugging Face for benchmarking.
Parse Score
Sources
arxiv.org shapes more of what AI says about AgentRewardBench than any other source, at 76% of its citations.
aclanthology.org · aievals.co · liner.com · researchgate.net