OLMo 2 32B Instruct March 2025 — benchmark results

Ai2's fully open OLMo 2 32B instruct model (Tulu 3.1 SFT+DPO+RLVR), billed as the first fully open model to beat GPT-3.5 Turbo and GPT-4o mini (March 2025). Provider: Allen AI. Released 2025-03-13. Access: Open.

Unified ELO 1452 ± 30, rank #989 of 1776 rated models, from 11 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Safety Anthropic Red Team99.3LM Evaluated Safety score (%)66.3
HELM Safety HarmBench84.1LM Evaluated Safety score (%)62.2
HELM Safety XSTest95.4LM Evaluated Safety score (%)44.8
HELM Safety SimpleSafetyTests98LM Evaluated Safety score (%)40.1
HELM Safety89.6Mean score (self-reported)37.2
HELM Capabilities - IFEval77.97IFEval Strict Acc26
HELM Capabilities - WildBench73.41WB Score22
HELM Capabilities - MMLU-Pro41.4COT correct16
HELM Capabilities - Omni-MATH16.07Acc16
HELM Safety BBQ71.4BBQ accuracy (%)10.5
HELM Capabilities - GPQA28.7COT correct6

Interactive version: theaggregate.ai/model?slug=olmo-2-32b-instruct-march-2025 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.