OLMo 2 32B Instruct March 2025 — benchmark results
Ai2's fully open OLMo 2 32B instruct model (Tulu 3.1 SFT+DPO+RLVR), billed as the first fully open model to beat GPT-3.5 Turbo and GPT-4o mini (March 2025). Provider: Allen AI. Released 2025-03-13. Access: Open.
Unified ELO 1452 ± 30, rank #989 of 1776 rated models, from 11 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Safety Anthropic Red Team | 99.3 | LM Evaluated Safety score (%) | 66.3 |
| HELM Safety HarmBench | 84.1 | LM Evaluated Safety score (%) | 62.2 |
| HELM Safety XSTest | 95.4 | LM Evaluated Safety score (%) | 44.8 |
| HELM Safety SimpleSafetyTests | 98 | LM Evaluated Safety score (%) | 40.1 |
| HELM Safety | 89.6 | Mean score (self-reported) | 37.2 |
| HELM Capabilities - IFEval | 77.97 | IFEval Strict Acc | 26 |
| HELM Capabilities - WildBench | 73.41 | WB Score | 22 |
| HELM Capabilities - MMLU-Pro | 41.4 | COT correct | 16 |
| HELM Capabilities - Omni-MATH | 16.07 | Acc | 16 |
| HELM Safety BBQ | 71.4 | BBQ accuracy (%) | 10.5 |
| HELM Capabilities - GPQA | 28.7 | COT correct | 6 |
Interactive version: theaggregate.ai/model?slug=olmo-2-32b-instruct-march-2025 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.