OLMo 3.1 32B (Thinking) — benchmark results
OLMo 3.1 32B evaluated with thinking enabled. Provider: Allen AI. Released 2026-01-15. Access: Open.
Unified ELO 1506 ± 12, rank #776 of 1776 rated models, from 365 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EuroEval Dutch NLU - CoNLL NL | 81.64 | Named entity recognition Score (%) | 100 |
| EuroEval Danish NLU - Angry Tweets | 59.97 | Sentiment classification Score (%) | 97.2 |
| EuroEval German NLU - GermEval | 74.37 | Named entity recognition Score (%) | 97.2 |
| EuroEval Catalan NLU - WikiANN CA | 75.06 | Named entity recognition Score (%) | 95.7 |
| Medmarks - Med-HALT Reasoning NOTA | 72.07 | Score (%) | 94.3 |
| EuroEval French NLU - Eltec | 72.07 | Named entity recognition Score (%) | 93.9 |
| EuroEval Finnish NLU - Turku NER FI | 71.69 | Named entity recognition Score (%) | 93.8 |
| EuroEval Portuguese NLU - HAREM | 60.41 | Named entity recognition Score (%) | 93.5 |
| YapBench | 563 | YapIndex (lower is better) | 93.2 |
| EuroEval Romanian Common Sense Reasoning | 58.01 | Common Sense Reasoning Average Score (%) | 93.1 |
| EuroEval English NLU - CoNLL EN | 84.82 | Named entity recognition Score (%) | 92.9 |
| EuroEval Danish NLU - Dansk | 67.47 | Named entity recognition Score (%) | 92.6 |
Interactive version: theaggregate.ai/model?slug=olmo-3-1-32b-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.