OLMo-2-1124-13B-Instruct: benchmark results
Provider: Allen AI. Released 2024-11-26. Access: Open.
Unified ELO 1428 ± 24, rank #1097 of 1639 rated models, from 145 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EuroEval English NLU - SST-5 | 69.45 | Sentiment classification Score (%) | 94.9 |
| HREF | 35.6 | Average HREF Score (%) | 78.8 |
| EuroEval English NLU - SQuAD | 83.61 | Reading comprehension Score (%) | 77.1 |
| EuroEval Portuguese NLU - MultiWikiQA PT | 72.88 | Reading comprehension Score (%) | 73.3 |
| Open PL LLM - PolEmo2-OUT (multiple choice, 5-shot) | 73.68 | Accuracy (%) | 72.9 |
| EuroEval French NLU - Allocine | 94.09 | Sentiment classification Score (%) | 72.1 |
| EuroEval English NLU | 67.97 | NLU Average Score (%) | 69.4 |
| EuroEval English | 68.01 | Average Score (%) | 67.2 |
| EuroEval English NLU - ScaLA EN | 48.67 | Linguistic acceptability Score (%) | 66.4 |
| EuroEval Dutch NLU - SQuAD NL | 74.61 | Reading comprehension Score (%) | 65.4 |
| EuroEval Italian NLU - SQuAD IT | 69.44 | Reading comprehension Score (%) | 65.3 |
| EuroEval Spanish NLU - Sentiment Headlines ES | 43.62 | Sentiment classification Score (%) | 64.1 |
Interactive version: theaggregate.ai/model?slug=olmo-2-1124-13b-instruct · How It Works · Data refreshed daily, snapshot 2026-10-09.