OLMo 3 32B (Thinking) — benchmark results

OLMo 3 32B evaluated with thinking enabled. Provider: Allen AI. Released 2025-11-20. Access: Open.

Unified ELO 1515 ± 31, rank #748 of 1776 rated models, from 73 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Medmarks - SCTpublic67.64Score (%)78.6
AA LiveCodeBench67.2Pass@1 (%)75.1
Medmarks - Med-HALT Reasoning NOTA63.53Score (%)72.9
AA AIME 202573.67Accuracy (%)69.1
Medmarks - Med-HALT Reasoning FCT71.67Score (%)65.7
Medmarks - PubHealthBench Reviewed84.04Score (%)58.6
AA IFBench49.12Accuracy (%)58.3
Medmarks - LongHealth Task 186.5Score (%)55.7
Medmarks - MetaMedQA65.36Score (%)55.7
Medmarks - MedXpertQA Understanding22.18Score (%)52.1
AA MMLU-Pro75.94Accuracy (%)51.5
Medmarks - SuperGPQA Medicine Hard34.25Score (%)50

Interactive version: theaggregate.ai/model?slug=olmo-3-32b-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.