GPT-4o (2024-08-06) — benchmark results

August 6, 2024 GPT-4o snapshot, tracked when sources report the dated API model. Provider: OpenAI. Released 2024-08-06. Access: API.

Unified ELO 1616 ± 12, rank #441 of 1776 rated models, from 200 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AutoLogi62.04Score (self-reported)100
HELMET (128K)64.8Average Score100
LiveBench Connections58Score100
LongProc94.8Accuracy @ 0.5K (%)100
Open Portuguese LLM - HateBR Offensive93.2F1 (%)100
OpenVLM Video - MMBench-Video43Normalized Score (%)100
ReliableMath - Prudence1.5Score (%)100
HELM (Stanford)92.76Mean Win Rate (%)98.9
VNTL Leaderboard74.97Accuracy (%)98.8
LiveBench Coding Completion55.26Score98.6
BenchBench96.53Aggregate Score (%)98.5
HELM NaturalQuestions (Closed)49.58F1 (%)97.8

Interactive version: theaggregate.ai/model?slug=gpt-4o-2024-08-06 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.