GPT-4o: benchmark results
OpenAI GPT-4o omni model for general text, vision, and multimodal tasks. Provider: OpenAI. Released 2024-05-13. Access: API.
Unified ELO 1590 ± 1, rank #290 of 1392 rated models, from 1500 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AGC-Bench - science_analogies | 2.5 | Dataset z-score | 100 |
| AIPatient Arena | 4.99 | Questioning Skills (QS) (self-reported) | 100 |
| ALARB | 46 | Facts+regulations correct (self-reported) | 100 |
| BlueBench - News Classification | 69.07 | Score (%) | 100 |
| BlueBench - QA Finance | 37 | Score (%) | 100 |
| CMMMU | 53.1 | Test Overall (%) | 100 |
| CMMMU (Validation) | 52.2 | Validation Overall (%) | 100 |
| Chain-of-Procedure | 62.5 | Zero-Shot (self-reported) | 100 |
| CyberMetric | 92.45 | Accuracy (%) | 100 |
| DITING - Zero Pronoun Translation | 5.56 | Score | 100 |
| Do LLMs Build World Models From Text? A Multil | 32.4 | L3 (viewpoint) (self-reported) | 100 |
| EmbodiedBench Manipulation | 28.9 | Avg Score (%) | 100 |
Interactive version: theaggregate.ai/model?slug=gpt-4o · How It Works · Data refreshed daily, snapshot 2026-09-05.