Data as of Oct 5, 2026Based on 11,419 AI responses
Reviewed by Dimitry Apollonsky ·
BFCL is a Berkeley-run leaderboard that assesses large language models on their ability to call external functions (tools) accurately using real-world data. It publishes periodic updates and documentation on evaluation methodology, datasets, and metrics across versions, including AST, multi-turn interactions, and holistic agentic evaluation. The project provides access to code and data for reproducibility and benchmarking.
<1%No change
of AI answers about Berkeley Function Calling Leaderboard and its rivals. Since Jul 5
The market map
LLM Observability and Evaluation PlatformsMentioned in
Since Jul 5
Rank
Where Berkeley Function Calling Leaderboard ranks in AI
gorilla.cs.berkeley.edu 0%Other sites 100%