OLMo-2-0325-32B-Instruct: benchmark results
Provider: Allen AI. Released 2025-03-13. Access: Open.
Unified ELO 1481 ± 17, rank #833 of 1639 rated models, from 303 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Grip on LLMs - Dutch Simplification - Amsterdam | 45.14 | SARI (0-100) | 96.7 |
| Grip on LLMs - Dutch Simplification Average | 43.71 | SARI (0-100; mean of two datasets) | 96.7 |
| EuroEval English NLU - SST-5 | 69.75 | Sentiment classification Score (%) | 95.8 |
| EuroEval Dutch NLU - DBRD | 92.51 | Sentiment classification Score (%) | 95 |
| EuroEval Greek NLU - Greek SA | 80.66 | Sentiment classification Score (%) | 92.6 |
| Open Portuguese LLM - FaQuAD NLI | 80.67 | Macro F1 (%) | 90.7 |
| EuroEval Serbian NLU - MMS SR | 49.82 | Sentiment classification Score (%) | 90.6 |
| EuroEval Norwegian NLU - Norec | 59.7 | Sentiment classification Score (%) | 90 |
| Grip on LLMs - Dutch Simplification - INT Duidelijke Taal | 42.29 | SARI (0-100) | 90 |
| Grip on LLMs - HonestCityBench - Average | 38 | Appropriate acknowledgement rate (%; mean of five categories | 90 |
| Grip on LLMs - HonestCityBench - Outdated Information | 53 | Appropriate acknowledgement rate (%; LLM judge) | 90 |
| EuroEval Swedish NLU - Swerec | 78.94 | Sentiment classification Score (%) | 89.8 |
Interactive version: theaggregate.ai/model?slug=olmo-2-0325-32b-instruct · How It Works · Data refreshed daily, snapshot 2026-10-09.