falcon-40B: benchmark results
Provider: TII. Released 2023-05-25. Access: Open.
Unified ELO 1393 ± 1, rank #1164 of 1392 rated models, from 52 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Classic - WikiFact | 38.03 | Exact Match (%) | 95.5 |
| HELM Classic - NaturalQuestions Closed Book | 39.22 | F1 (%) | 92.4 |
| HELM Classic - MATH | 20.98 | Equivalent (%) | 88.2 |
| HELM Classic - IMDB | 95.9 | Exact Match (%) | 87.9 |
| HELM Classic - MMLU | 50.89 | Exact Match (%) | 87.9 |
| HELM Classic - NaturalQuestions Open Book | 67.53 | F1 (%) | 86.2 |
| HELM Classic - bAbI | 57.62 | Exact Match (%) | 85.5 |
| HellaSwag | 85.28 | Accuracy (%) | 85.5 |
| HELM Classic - LegalSupport | 60.53 | Exact Match (%) | 85.3 |
| HELM Classic - GSM8K | 25 | Exact Match (%) | 82.4 |
| HELM Classic - MATH Chain-of-Thought | 13.01 | Equivalent (%) | 82.4 |
| HELM Classic - Entity Data Imputation | 83.01 | Exact Match (%) | 81.8 |
Interactive version: theaggregate.ai/model?slug=falcon-40b · How It Works · Data refreshed daily, snapshot 2026-09-05.