OLMoE 1B-7B Instruct January 2025 — benchmark results
Provider: Allen AI. Released 2025-01-27. Access: Open.
Unified ELO 1246 ± 35, rank #1627 of 1776 rated models, from 11 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Safety HarmBench | 62.9 | LM Evaluated Safety score (%) | 29.1 |
| HELM Safety XSTest | 84.3 | LM Evaluated Safety score (%) | 4.7 |
| HELM Capabilities - IFEval | 62.85 | IFEval Strict Acc | 4 |
| HELM Safety Anthropic Red Team | 88.9 | LM Evaluated Safety score (%) | 3.5 |
| HELM Safety | 70.1 | Mean score (self-reported) | 2.5 |
| HELM Safety SimpleSafetyTests | 72.5 | LM Evaluated Safety score (%) | 2.3 |
| HELM Capabilities - GPQA | 21.97 | COT correct | 2 |
| HELM Capabilities - Omni-MATH | 9.33 | Acc | 2 |
| HELM Capabilities - WildBench | 55.11 | WB Score | 2 |
| HELM Capabilities - MMLU-Pro | 16.9 | COT correct | 0 |
| HELM Safety BBQ | 41.9 | BBQ accuracy (%) | 0 |
Interactive version: theaggregate.ai/model?slug=olmoe-1b-7b-instruct-january-2025 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.