BFCL V4 - Irrelevance Detection: leaderboard

Metric: Accuracy (%). Source: gorilla.cs.berkeley.edu. 109 models tracked.

Top models

#ModelScore
1Ministral-8B-Instruct-2410100
2Llama 3.1 Nemotron Ultra 253B v1100
3Claude Haiku 4.5 (20251001)95.29
4Claude Sonnet 4.595.03
5Gemini 2.5 Flash93.67
6Gemini 2.5 Flash Lite93.33
7DeepSeek V3.2 Exp93.18
8Mistral Medium 391.95
9GPT-5 Mini91.01
10Claude Opus 4.5 (20251101)90.75
11GPT-5 Nano89.1
12Mistral Small 3.287.94
13Phi-487.55
14Kimi K287.34
15Falcon3-1B-Instruct87.3

Interactive version: theaggregate.ai/benchmark?slug=bfcl-v4-irrelevance-detection · How It Works · Data refreshed daily, snapshot 2026-09-19.