BFCL v2 — leaderboard
Berkeley Function Calling Leaderboard (BFCL) v2 is a comprehensive benchmark for evaluating large language models' function calling capabilities. It features 2,251 question-function-answer pairs with enterprise and OSS-contributed functions, addressing data contamination and bias through live, user-contributed scenarios. The benchmark evaluates AST accuracy, executable accuracy, irrelevance detection, and relevance detection across multiple programming languages (Python, Java, JavaScript) and includes complex real-world function calling scenarios with multi-lingual prompts.
Source: gorilla.cs.berkeley.edu.
Interactive version: theaggregate.ai/benchmark?slug=bfcl-v2 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.