BFCL V4 - Irrelevance Detection: leaderboard
Metric: Accuracy (%). Source: gorilla.cs.berkeley.edu. 109 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Ministral-8B-Instruct-2410 | 100 |
| 2 | Llama 3.1 Nemotron Ultra 253B v1 | 100 |
| 3 | Claude Haiku 4.5 (20251001) | 95.29 |
| 4 | Claude Sonnet 4.5 | 95.03 |
| 5 | Gemini 2.5 Flash | 93.67 |
| 6 | Gemini 2.5 Flash Lite | 93.33 |
| 7 | DeepSeek V3.2 Exp | 93.18 |
| 8 | Mistral Medium 3 | 91.95 |
| 9 | GPT-5 Mini | 91.01 |
| 10 | Claude Opus 4.5 (20251101) | 90.75 |
| 11 | GPT-5 Nano | 89.1 |
| 12 | Mistral Small 3.2 | 87.94 |
| 13 | Phi-4 | 87.55 |
| 14 | Kimi K2 | 87.34 |
| 15 | Falcon3-1B-Instruct | 87.3 |
Interactive version: theaggregate.ai/benchmark?slug=bfcl-v4-irrelevance-detection · How It Works · Data refreshed daily, snapshot 2026-09-19.