Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
SkillsBench is an evaluation framework that measures how well AI agent skills perform across diverse, expert-curated tasks in high-GDP domains. It benchmarks agents across three abstraction layers—skills, agent harness, and models—to quantify skill effectiveness and model capability.
Parse Score
Sources
playbooks.com shapes more of what AI says about SkillsBench than any other source, at 100% of its citations.