Falcon3-3B-Instruct: benchmark results
Provider: TII. Released 2024-12-17. Access: Open.
Unified ELO 1343 ± 20, rank #2270 of 2656 rated models, from 107 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open LLM Leaderboard - IFEval | 69.77 | Score | 82.8 |
| MEDIC Benchmark | 75.28 | MEDIC Public Table Average | 82.1 |
| Open LLM Leaderboard - MATH Level 5 | 25 | Score | 78.1 |
| EuroEval Portuguese NLU - HAREM | 44.97 | Named entity recognition Score (%) | 61.8 |
| Open Portuguese LLM - FaQuAD NLI | 70.01 | Macro F1 (%) | 60.2 |
| EuroEval Portuguese NLU - MultiWikiQA PT | 70.43 | Reading comprehension Score (%) | 59.4 |
| EuroEval Portuguese NLU - ScaLA PT | 14.58 | Linguistic acceptability Score (%) | 53.9 |
| Open LLM Leaderboard - MuSR | 41.36 | Score | 53.6 |
| BFCL V4 - Relevance Detection | 81.25 | Accuracy (%) | 50 |
| EuroEval Portuguese NLU | 49.39 | NLU Average Score (%) | 49.2 |
| EuroEval Portuguese | 41.48 | Average Score (%) | 46.1 |
| Open LLM Leaderboard - GPQA | 28.86 | Score | 43.2 |
Interactive version: theaggregate.ai/model?slug=falcon3-3b-instruct · How It Works · Data refreshed daily, snapshot 2026-09-19.