GPT-4 Base: benchmark results
Provider: OpenAI. Released 2023-03-14. Access: API.
Unified ELO 1499 ± 47, rank #1107 of 2656 rated models, from 23 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SAD - Anti-Imitation | 16.82 | Score (%, higher is better) | 85 |
| SAD - Anti-Imitation (Situating Prompt) | 16.82 | Score (%, higher is better) | 85 |
| SAD - Introspection (Situating Prompt) | 35.04 | Score (%, higher is better) | 80 |
| SAD - Stages | 47.31 | Score (%, higher is better) | 80 |
| PubMedQA | 80.4 | Accuracy (%) | 78.3 |
| SAD - Introspection | 34.42 | Score (%, higher is better) | 75 |
| SAD - Predict Words | 35.65 | Self-prediction score (%, higher is better) | 68.2 |
| SAD - Predict Words (Situating Prompt) | 34.35 | Self-prediction score (%, higher is better) | 68.2 |
| SAD - Stages (Situating Prompt) | 43.31 | Score (%, higher is better) | 55 |
| SAD - Facts (Situating Prompt) | 62.12 | Score (%, higher is better) | 50 |
| SAD - Influence | 62.19 | Score (%, higher is better) | 50 |
| SAD - Mini | 57.16 | Score (%, higher is better) | 50 |
Interactive version: theaggregate.ai/model?slug=gpt-4-base · How It Works · Data refreshed daily, snapshot 2026-09-19.