OLMo 3.1 32B (Thinking) — benchmark results

OLMo 3.1 32B evaluated with thinking enabled. Provider: Allen AI. Released 2026-01-15. Access: Open.

Unified ELO 1506 ± 12, rank #776 of 1776 rated models, from 365 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EuroEval Dutch NLU - CoNLL NL81.64Named entity recognition Score (%)100
EuroEval Danish NLU - Angry Tweets59.97Sentiment classification Score (%)97.2
EuroEval German NLU - GermEval74.37Named entity recognition Score (%)97.2
EuroEval Catalan NLU - WikiANN CA75.06Named entity recognition Score (%)95.7
Medmarks - Med-HALT Reasoning NOTA72.07Score (%)94.3
EuroEval French NLU - Eltec72.07Named entity recognition Score (%)93.9
EuroEval Finnish NLU - Turku NER FI71.69Named entity recognition Score (%)93.8
EuroEval Portuguese NLU - HAREM60.41Named entity recognition Score (%)93.5
YapBench563YapIndex (lower is better)93.2
EuroEval Romanian Common Sense Reasoning58.01Common Sense Reasoning Average Score (%)93.1
EuroEval English NLU - CoNLL EN84.82Named entity recognition Score (%)92.9
EuroEval Danish NLU - Dansk67.47Named entity recognition Score (%)92.6

Interactive version: theaggregate.ai/model?slug=olmo-3-1-32b-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.