GPT-4o: benchmark results

OpenAI GPT-4o omni model for general text, vision, and multimodal tasks. Provider: OpenAI. Released 2024-05-13. Access: API.

Unified ELO 1590 ± 1, rank #290 of 1392 rated models, from 1500 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AGC-Bench - science_analogies2.5Dataset z-score100
AIPatient Arena4.99Questioning Skills (QS) (self-reported)100
ALARB46Facts+regulations correct (self-reported)100
BlueBench - News Classification69.07Score (%)100
BlueBench - QA Finance37Score (%)100
CMMMU53.1Test Overall (%)100
CMMMU (Validation)52.2Validation Overall (%)100
Chain-of-Procedure62.5Zero-Shot (self-reported)100
CyberMetric92.45Accuracy (%)100
DITING - Zero Pronoun Translation5.56Score100
Do LLMs Build World Models From Text? A Multil32.4L3 (viewpoint) (self-reported)100
EmbodiedBench Manipulation28.9Avg Score (%)100

Interactive version: theaggregate.ai/model?slug=gpt-4o · How It Works · Data refreshed daily, snapshot 2026-09-05.