GPT-4o (Mar 2025): benchmark results
March 2025 GPT-4o snapshot kept separate from other GPT-4o rows because benchmark sources report distinct scores. Provider: OpenAI. Released 2025-03-01. Access: API.
Unified ELO 1545 ± 1, rank #654 of 1761 rated models, from 9 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LLM Emergent Collusion | 23 | Collusion Rate (%) | 91.7 |
| Elimination Game (Lechmazur) | 5.5 | TrueSkill μ | 89.8 |
| Generalization V1 (Lechmazur) | 1.97 | Avg Rank (lower is better) | 43.1 |
| NYT Connections Older Models | 11.8 | Score (%) | 42.7 |
| AA GPQA Diamond | 65.45 | Accuracy (%) | 40.8 |
| Step Game (Lechmazur) | 1.55 | TrueSkill μ | 39.2 |
| Artificial Analysis Intelligence Index | 6.5 | Intelligence Index | 39 |
| Confabulation Leaderboard (Lechmazur) | 38.12 | Confabulation rate % (lower is better) | 16.7 |
| AA Humanity's Last Exam | 3.97 | Accuracy (%) | 12.6 |
Interactive version: theaggregate.ai/model?slug=gpt-4o-mar-2025 · How It Works · Data refreshed daily, snapshot 2026-09-05.