GPT-4o (Mar 2025) — benchmark results

March 2025 GPT-4o snapshot kept separate from other GPT-4o rows because benchmark sources report distinct scores. Provider: OpenAI. Released 2025-03-01. Access: API.

Unified ELO 1561 ± 33, rank #587 of 1776 rated models, from 14 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Elimination Game (Lechmazur)5.5TrueSkill μ89.8
AA MMLU-Pro80.29Accuracy (%)68
AA MATH-50089.27Accuracy (%)63.6
AA SciCode36.57Accuracy (%)60.1
AA LiveCodeBench42.54Pass@1 (%)50.4
NYT Connections Older Models24.5Score (%)46.3
AA GPQA Diamond65.45Accuracy (%)44.3
Generalization V1 (Lechmazur)1.97Avg Rank (lower is better)43.1
Artificial Analysis Intelligence Index12.31Intelligence Index41.9
Step Game (Lechmazur)1.55TrueSkill μ39.2
AA Humanity's Last Exam4.96Accuracy (%)32.4
AA AIME 202525.67Accuracy (%)27.2

Interactive version: theaggregate.ai/model?slug=gpt-4o-mar-2025 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.