GPT-4o — benchmark results

OpenAI GPT-4o omni model for general text, vision, and multimodal tasks. Provider: OpenAI. Released 2024-05-13. Access: API.

Unified ELO 1643 ± 6, rank #386 of 1776 rated models, from 1474 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AGC-Bench - science_analogies2.5Dataset z-score100
AIPatient Arena4.99Questioning Skills (QS) (self-reported)100
ALARB46Facts+regulations correct (self-reported)100
BlueBench - News Classification69.07Score (%)100
BlueBench - QA Finance37Score (%)100
CMMMU53.1Test Overall (%)100
CMMMU (Validation)52.2Validation Overall (%)100
Chain-of-Procedure62.5Zero-Shot (self-reported)100
CyberMetric92.45Accuracy (%)100
DITING - Zero Pronoun Translation5.56Score100
Do LLMs Build World Models From Text? A Multil32.4L3 (viewpoint) (self-reported)100
FinBen (Financial LLM)46.01Average Score100

Interactive version: theaggregate.ai/model?slug=gpt-4o · How the rankings work · Data refreshed daily, snapshot 2026-07-22.