Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
Sage is a Princeton research group that develops open world evaluations and benchmarks for frontier AI agents, focusing on long-horizon, real-world tasks. They create standardized leaderboards and methodologies to assess agent reliability, computational reproducibility, and practical utility beyond traditional benchmarks.
Parse Score
Sources
arxiv.org shapes more of what AI says about SAgE (Science of Agent Evaluation) than any other source, at 50% of its citations.
sage.cs.princeton.edu