Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
TRAJECT-Bench is a comprehensive benchmark for evaluating tool-using language models across multiple practical domains. It tests models' ability to effectively utilize external tools for real-world tasks through parallel and sequential tool calling trajectories.
Parse Score
Sources
arxiv.org shapes more of what AI says about TRAJECT-Bench than any other source, at 80% of its citations.
openreview.net