Olmo 3.1 32B Instruct — benchmark results
Ai2's 32B instruct model from the fully open Olmo 3.1 update, scaling the Olmo 3 SFT+DPO+RLVR chat recipe up from 7B (December 2025). Provider: Allen AI. Released 2025-12-23. Access: Open.
Unified ELO 1506 ± 10, rank #775 of 1776 rated models, from 406 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AGC-Bench - story_generation_rocstories | 3.29 | Dataset z-score | 98.8 |
| EuroEval Catalan NLU - MultiWikiQA CA | 77.23 | Reading comprehension Score (%) | 97.2 |
| AGC-Bench - proparalogy | 0.87 | Dataset z-score | 96.3 |
| EuroEval Dutch Summarization - Wiki Lingua NL | 34.91 | Score (%) | 91.3 |
| EuroEval English NLU - ScaLA EN | 60.51 | Linguistic acceptability Score (%) | 90.4 |
| EuroEval Portuguese NLU - MultiWikiQA PT | 76.18 | Reading comprehension Score (%) | 89.6 |
| AGC-Bench - story_quality | 1.11 | Dataset z-score | 89.5 |
| EuroEval Portuguese NLU | 60.63 | NLU Average Score (%) | 89.1 |
| EuroEval Portuguese NLU - ScaLA PT | 38.51 | Linguistic acceptability Score (%) | 88.9 |
| EuroEval Dutch NLU - DBRD | 91.44 | Sentiment classification Score (%) | 87.1 |
| EuroEval Italian NLU - ScaLA IT | 41.26 | Linguistic acceptability Score (%) | 87.1 |
| EuroEval French NLU - ScaLA FR | 50.41 | Linguistic acceptability Score (%) | 87 |
Interactive version: theaggregate.ai/model?slug=olmo-3-1-32b-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.