Data as of Aug 25, 2026 · Based on 3,181,687 AI responses across 10,525 prompts · See how Parse measures this
ToolBench is a benchmark for evaluating large language models on software tool manipulation tasks, featuring diverse real-world tools like OpenWeather and Google Sheets. It provides infrastructure to measure execution success rates and supports both OpenAI and HuggingFace models for comparison.
Parse Score