Olmo 3.1 32B Instruct — benchmark results

Ai2's 32B instruct model from the fully open Olmo 3.1 update, scaling the Olmo 3 SFT+DPO+RLVR chat recipe up from 7B (December 2025). Provider: Allen AI. Released 2025-12-23. Access: Open.

Unified ELO 1506 ± 10, rank #775 of 1776 rated models, from 406 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AGC-Bench - story_generation_rocstories3.29Dataset z-score98.8
EuroEval Catalan NLU - MultiWikiQA CA77.23Reading comprehension Score (%)97.2
AGC-Bench - proparalogy0.87Dataset z-score96.3
EuroEval Dutch Summarization - Wiki Lingua NL34.91Score (%)91.3
EuroEval English NLU - ScaLA EN60.51Linguistic acceptability Score (%)90.4
EuroEval Portuguese NLU - MultiWikiQA PT76.18Reading comprehension Score (%)89.6
AGC-Bench - story_quality1.11Dataset z-score89.5
EuroEval Portuguese NLU60.63NLU Average Score (%)89.1
EuroEval Portuguese NLU - ScaLA PT38.51Linguistic acceptability Score (%)88.9
EuroEval Dutch NLU - DBRD91.44Sentiment classification Score (%)87.1
EuroEval Italian NLU - ScaLA IT41.26Linguistic acceptability Score (%)87.1
EuroEval French NLU - ScaLA FR50.41Linguistic acceptability Score (%)87

Interactive version: theaggregate.ai/model?slug=olmo-3-1-32b-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.