GPT-4o (Mar 2025): benchmark results

March 2025 GPT-4o snapshot kept separate from other GPT-4o rows because benchmark sources report distinct scores. Provider: OpenAI. Released 2025-03-01. Access: API.

Unified ELO 1545 ± 1, rank #654 of 1761 rated models, from 9 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LLM Emergent Collusion23Collusion Rate (%)91.7
Elimination Game (Lechmazur)5.5TrueSkill μ89.8
Generalization V1 (Lechmazur)1.97Avg Rank (lower is better)43.1
NYT Connections Older Models11.8Score (%)42.7
AA GPQA Diamond65.45Accuracy (%)40.8
Step Game (Lechmazur)1.55TrueSkill μ39.2
Artificial Analysis Intelligence Index6.5Intelligence Index39
Confabulation Leaderboard (Lechmazur)38.12Confabulation rate % (lower is better)16.7
AA Humanity's Last Exam3.97Accuracy (%)12.6

Interactive version: theaggregate.ai/model?slug=gpt-4o-mar-2025 · How It Works · Data refreshed daily, snapshot 2026-09-05.