OLMo-2-0325-32B-Instruct: benchmark results

Provider: Allen AI. Released 2025-03-13. Access: Open.

Unified ELO 1481 ± 17, rank #833 of 1639 rated models, from 303 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Grip on LLMs - Dutch Simplification - Amsterdam45.14SARI (0-100)96.7
Grip on LLMs - Dutch Simplification Average43.71SARI (0-100; mean of two datasets)96.7
EuroEval English NLU - SST-569.75Sentiment classification Score (%)95.8
EuroEval Dutch NLU - DBRD92.51Sentiment classification Score (%)95
EuroEval Greek NLU - Greek SA80.66Sentiment classification Score (%)92.6
Open Portuguese LLM - FaQuAD NLI80.67Macro F1 (%)90.7
EuroEval Serbian NLU - MMS SR49.82Sentiment classification Score (%)90.6
EuroEval Norwegian NLU - Norec59.7Sentiment classification Score (%)90
Grip on LLMs - Dutch Simplification - INT Duidelijke Taal42.29SARI (0-100)90
Grip on LLMs - HonestCityBench - Average38Appropriate acknowledgement rate (%; mean of five categories90
Grip on LLMs - HonestCityBench - Outdated Information53Appropriate acknowledgement rate (%; LLM judge)90
EuroEval Swedish NLU - Swerec78.94Sentiment classification Score (%)89.8

Interactive version: theaggregate.ai/model?slug=olmo-2-0325-32b-instruct · How It Works · Data refreshed daily, snapshot 2026-10-09.