falcon-11B — benchmark results
TII's 11B Falcon 2 generation model (May 2024), trained on 5.5T tokens with multilingual support, released under the permissive TII Falcon License 2.0. Provider: TII. Released 2024-05-14. Access: Open.
Unified ELO 1420 ± 18, rank #1127 of 1776 rated models, from 46 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EuroEval Portuguese NLU - SST-2 PT | 83.44 | Sentiment classification Score (%) | 84.3 |
| HellaSwag | 82.91 | Accuracy (%) | 76.3 |
| EuroEval Dutch NLU - DBRD | 90.49 | Sentiment classification Score (%) | 75.8 |
| EuroEval Italian NLU - Sentipolc16 | 59.88 | Sentiment classification Score (%) | 73.7 |
| WinoGrande | 78.3 | Accuracy (%) | 70 |
| Open PL LLM - RAG | 65.34 | Average RAG Score (%) | 66.4 |
| EuroEval Spanish NLU - ScaLA ES | 25.29 | Linguistic acceptability Score (%) | 65.2 |
| EuroEval Spanish NLU - Sentiment Headlines ES | 43.53 | Sentiment classification Score (%) | 63.6 |
| EuroEval Portuguese NLU - ScaLA PT | 16.64 | Linguistic acceptability Score (%) | 59.3 |
| Open PL LLM - Generative | 50.78 | Average Generative Score (%) | 54.4 |
| GSM8K | 53.83 | Accuracy (%) | 54.3 |
| EuroEval Portuguese NLU | 50.52 | NLU Average Score (%) | 53.7 |
Interactive version: theaggregate.ai/model?slug=falcon-11b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.