Falcon3-3B-Base: benchmark results
Provider: TII. Access: Open.
Unified ELO 1311 ± 28, rank #2415 of 2656 rated models, from 92 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Swallow - Pre-trained English - GSM8K | 63.4 | Exact Match (%) | 57.1 |
| Swallow - Pre-trained English - MATH | 34.4 | Exact Match (%) | 57.1 |
| Open LLM Leaderboard - MATH Level 5 | 11.78 | Score | 53.1 |
| Open LLM Leaderboard - GPQA | 29.7 | Score | 52.6 |
| Swallow - Pre-trained English - BBH | 54.7 | Exact Match (%) | 51 |
| Swallow - Pre-trained English - HumanEval | 34.8 | Pass@1 (%) | 46.9 |
| EuroEval Portuguese NLU - HAREM | 40.03 | Named entity recognition Score (%) | 45.2 |
| Swallow - Pre-trained English - Average | 49.4 | Average Score (%) | 44.9 |
| EuroEval Portuguese NLU - SST-2 PT | 73.48 | Sentiment classification Score (%) | 43.2 |
| Open Portuguese LLM - FaQuAD NLI | 58.64 | Macro F1 (%) | 42.6 |
| Open Portuguese LLM - OAB Exams | 39.36 | Accuracy (%) | 41.1 |
| Swallow - Pre-trained English - SQuAD2 | 50.3 | Exact Match (%) | 39.8 |
Interactive version: theaggregate.ai/model?slug=falcon3-3b-base · How It Works · Data refreshed daily, snapshot 2026-09-19.