GPT-4o (0513): benchmark results
Provider: OpenAI. Released 2024-05-13. Access: API.
Unified ELO 1720 ± 28, rank #288 of 2656 rated models, from 38 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LongVideoBench | 66.7 | Test Total Accuracy (%) | 100 |
| MM-MT-Bench | 7.72 | Score (points) | 100 |
| SeaEval | 72.86 | Average Score (%) | 100 |
| SeaEval - Cultural Reasoning - PH-Eval (Zero-Shot) | 77 | Accuracy (%) | 100 |
| SeaEval - Cultural Reasoning - SG-Eval (Zero-Shot) | 84.47 | Accuracy (%) | 100 |
| SeaEval - Cultural Reasoning - SG-Eval v1 Cleaned (Zero-Shot) | 80.88 | Accuracy (%) | 100 |
| SeaEval - Cultural Reasoning - SG-Eval v2 Open (Zero-Shot) | 57.28 | Accuracy (%) | 100 |
| SeaEval - FLORES Translation - Malay-to-English (Zero-Shot) | 45.15 | BLEU (0-100) | 100 |
| SeaEval - Fundamental NLP Tasks - C3 (Zero-Shot) | 96.48 | Accuracy (%) | 100 |
| SeaEval - Fundamental NLP Tasks - QNLI (Zero-Shot) | 93.04 | Accuracy (%) | 100 |
| SeaEval - Fundamental NLP Tasks - WNLI (Zero-Shot) | 92.96 | Accuracy (%) | 100 |
| SeaEval - Multilingual Reasoning - IndoMMLU (Zero-Shot) | 75.85 | Accuracy (%) | 100 |
Interactive version: theaggregate.ai/model?slug=gpt-4o-0513 · How It Works · Data refreshed daily, snapshot 2026-09-19.