GPT-4o (0513): benchmark results

Provider: OpenAI. Released 2024-05-13. Access: API.

Unified ELO 1720 ± 28, rank #288 of 2656 rated models, from 38 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LongVideoBench66.7Test Total Accuracy (%)100
MM-MT-Bench7.72Score (points)100
SeaEval72.86Average Score (%)100
SeaEval - Cultural Reasoning - PH-Eval (Zero-Shot)77Accuracy (%)100
SeaEval - Cultural Reasoning - SG-Eval (Zero-Shot)84.47Accuracy (%)100
SeaEval - Cultural Reasoning - SG-Eval v1 Cleaned (Zero-Shot)80.88Accuracy (%)100
SeaEval - Cultural Reasoning - SG-Eval v2 Open (Zero-Shot)57.28Accuracy (%)100
SeaEval - FLORES Translation - Malay-to-English (Zero-Shot)45.15BLEU (0-100)100
SeaEval - Fundamental NLP Tasks - C3 (Zero-Shot)96.48Accuracy (%)100
SeaEval - Fundamental NLP Tasks - QNLI (Zero-Shot)93.04Accuracy (%)100
SeaEval - Fundamental NLP Tasks - WNLI (Zero-Shot)92.96Accuracy (%)100
SeaEval - Multilingual Reasoning - IndoMMLU (Zero-Shot)75.85Accuracy (%)100

Interactive version: theaggregate.ai/model?slug=gpt-4o-0513 · How It Works · Data refreshed daily, snapshot 2026-09-19.