GPT-4 Base: benchmark results

Provider: OpenAI. Released 2023-03-14. Access: API.

Unified ELO 1499 ± 47, rank #1107 of 2656 rated models, from 23 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
SAD - Anti-Imitation16.82Score (%, higher is better)85
SAD - Anti-Imitation (Situating Prompt)16.82Score (%, higher is better)85
SAD - Introspection (Situating Prompt)35.04Score (%, higher is better)80
SAD - Stages47.31Score (%, higher is better)80
PubMedQA80.4Accuracy (%)78.3
SAD - Introspection34.42Score (%, higher is better)75
SAD - Predict Words35.65Self-prediction score (%, higher is better)68.2
SAD - Predict Words (Situating Prompt)34.35Self-prediction score (%, higher is better)68.2
SAD - Stages (Situating Prompt)43.31Score (%, higher is better)55
SAD - Facts (Situating Prompt)62.12Score (%, higher is better)50
SAD - Influence62.19Score (%, higher is better)50
SAD - Mini57.16Score (%, higher is better)50

Interactive version: theaggregate.ai/model?slug=gpt-4-base · How It Works · Data refreshed daily, snapshot 2026-09-19.