Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
τ bench is a benchmark that simulates dynamic conversations between a language model-powered user and a language agent equipped with domain-specific API tools and policy guidelines. It evaluates tool-agent user interactions in real-world domains such as airline and retail tasks.
Parse Score