GPT-4o (Mar 2025) — benchmark results
March 2025 GPT-4o snapshot kept separate from other GPT-4o rows because benchmark sources report distinct scores. Provider: OpenAI. Released 2025-03-01. Access: API.
Unified ELO 1561 ± 33, rank #587 of 1776 rated models, from 14 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Elimination Game (Lechmazur) | 5.5 | TrueSkill μ | 89.8 |
| AA MMLU-Pro | 80.29 | Accuracy (%) | 68 |
| AA MATH-500 | 89.27 | Accuracy (%) | 63.6 |
| AA SciCode | 36.57 | Accuracy (%) | 60.1 |
| AA LiveCodeBench | 42.54 | Pass@1 (%) | 50.4 |
| NYT Connections Older Models | 24.5 | Score (%) | 46.3 |
| AA GPQA Diamond | 65.45 | Accuracy (%) | 44.3 |
| Generalization V1 (Lechmazur) | 1.97 | Avg Rank (lower is better) | 43.1 |
| Artificial Analysis Intelligence Index | 12.31 | Intelligence Index | 41.9 |
| Step Game (Lechmazur) | 1.55 | TrueSkill μ | 39.2 |
| AA Humanity's Last Exam | 4.96 | Accuracy (%) | 32.4 |
| AA AIME 2025 | 25.67 | Accuracy (%) | 27.2 |
Interactive version: theaggregate.ai/model?slug=gpt-4o-mar-2025 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.