Data as of Sep 29, 2026 · Based on 1,788 AI responses · See how Parse measures this
τ-bench is a benchmarking platform from Sierra for evaluating AI agents in collaborative, real-world scenarios, emphasizing coordination and tool use to achieve shared objectives across enterprise domains. It features a public leaderboard, results submission, and domain-specific tracks (retail, airline, telecom), with recent expansion into telecom and dual-control environments using user simulators. The project has notable milestones, including GPT-5 achieving state-of-the-art performance on τ-bench (96% telecom, 82% retail, 63% airline) and an ICLR 2025 paper acceptance, with ongoing updates and research.
<1%No change
of AI answers about τ‑Bench and its rivals. Week of Sep 21
AI mentioned τ‑Bench in <1% of answers about τ‑Bench and its rivals in the week of Sep 21.