GPT-4 (0613): benchmark results

June 13, 2023 GPT-4 API snapshot, kept separate from other GPT-4 rows when sources report it explicitly. Provider: OpenAI. Released 2023-06-13. Access: API.

Unified ELO 1537 ± 1, rank #501 of 1392 rated models, from 80 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
CommonGen45.11Overall (%)100
FastEval77.78Total Score100
HELM v2 Lite - HumanEval (Code)73Pass@1 (%)100
HELM v2 Lite - LegalBench71.01Exact Match (%)100
HELM v2 Lite - MedQA88.46Exact Match (%)100
InfiBench70.64Score (%)100
LogicKor - Reasoning9.42Score (0-10)98.3
MAGI-Hard77.85Accuracy (%, 2024-05 snapshot)98.3
HELM Lite90.89Mean win rate (self-reported)97.4
JustEval - Clarity4.99Score (1-5)96.7
JustEval - Factuality4.9Score (1-5)96.7
LogicKor - Grammar8.35Score (0-10)96.6

Interactive version: theaggregate.ai/model?slug=gpt-4-0613 · How It Works · Data refreshed daily, snapshot 2026-09-05.