GPT-4 — benchmark results

OpenAI's GPT-4 flagship model (March 2023). Provider: OpenAI. Released 2023-03-14. Access: API.

Unified ELO 1580 ± 19, rank #531 of 1776 rated models, from 160 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AgentBoard70Progress Rate (self-reported)100
BELLS91.7BELLS Score (%)100
ClassEval37.6Class-Level Pass@1 (%)100
CodeScope39.2Score100
CodeScope - Code Generation31.47Score100
CodeScope - Code Repair30.03Score100
CodeScope - Code Summarization33.66Score100
CodeScope - Code Translation31.29Score100
CodeScope - Program Synthesis36.36Score100
CyberBench (NLP)74.03Avg Score (%)100
EvoEval Combine53Pass@1 (%)100
EvoEval Concise81.1Pass@1 (%)100

Interactive version: theaggregate.ai/model?slug=gpt-4 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.