OLMo 3 32B (Thinking) — benchmark results
OLMo 3 32B evaluated with thinking enabled. Provider: Allen AI. Released 2025-11-20. Access: Open.
Unified ELO 1515 ± 31, rank #748 of 1776 rated models, from 73 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Medmarks - SCTpublic | 67.64 | Score (%) | 78.6 |
| AA LiveCodeBench | 67.2 | Pass@1 (%) | 75.1 |
| Medmarks - Med-HALT Reasoning NOTA | 63.53 | Score (%) | 72.9 |
| AA AIME 2025 | 73.67 | Accuracy (%) | 69.1 |
| Medmarks - Med-HALT Reasoning FCT | 71.67 | Score (%) | 65.7 |
| Medmarks - PubHealthBench Reviewed | 84.04 | Score (%) | 58.6 |
| AA IFBench | 49.12 | Accuracy (%) | 58.3 |
| Medmarks - LongHealth Task 1 | 86.5 | Score (%) | 55.7 |
| Medmarks - MetaMedQA | 65.36 | Score (%) | 55.7 |
| Medmarks - MedXpertQA Understanding | 22.18 | Score (%) | 52.1 |
| AA MMLU-Pro | 75.94 | Accuracy (%) | 51.5 |
| Medmarks - SuperGPQA Medicine Hard | 34.25 | Score (%) | 50 |
Interactive version: theaggregate.ai/model?slug=olmo-3-32b-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.