GPT-4o — benchmark results
OpenAI GPT-4o omni model for general text, vision, and multimodal tasks. Provider: OpenAI. Released 2024-05-13. Access: API.
Unified ELO 1643 ± 6, rank #386 of 1776 rated models, from 1474 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AGC-Bench - science_analogies | 2.5 | Dataset z-score | 100 |
| AIPatient Arena | 4.99 | Questioning Skills (QS) (self-reported) | 100 |
| ALARB | 46 | Facts+regulations correct (self-reported) | 100 |
| BlueBench - News Classification | 69.07 | Score (%) | 100 |
| BlueBench - QA Finance | 37 | Score (%) | 100 |
| CMMMU | 53.1 | Test Overall (%) | 100 |
| CMMMU (Validation) | 52.2 | Validation Overall (%) | 100 |
| Chain-of-Procedure | 62.5 | Zero-Shot (self-reported) | 100 |
| CyberMetric | 92.45 | Accuracy (%) | 100 |
| DITING - Zero Pronoun Translation | 5.56 | Score | 100 |
| Do LLMs Build World Models From Text? A Multil | 32.4 | L3 (viewpoint) (self-reported) | 100 |
| FinBen (Financial LLM) | 46.01 | Average Score | 100 |
Interactive version: theaggregate.ai/model?slug=gpt-4o · How the rankings work · Data refreshed daily, snapshot 2026-07-22.