BFCL v2: leaderboard

Berkeley Function Calling Leaderboard (BFCL) v2 is a benchmark for evaluating large language models' function calling capabilities. It features 2,251 question-function-answer pairs with enterprise and OSS-contributed functions, addressing data contamination and bias through live, user-contributed scenarios. The benchmark evaluates AST accuracy, executable accuracy, irrelevance detection, and relevance detection across Python, Java, and JavaScript and includes complex function calling scenarios with multi-lingual prompts.

Source: gorilla.cs.berkeley.edu.

Interactive version: theaggregate.ai/benchmark?slug=bfcl-v2 · How It Works · Data refreshed daily, snapshot 2026-09-05.