GPT-4 (0314) — benchmark results
March 14, 2023 GPT-4 API snapshot, kept separate from later GPT-4 snapshots. Provider: OpenAI. Released 2023-03-14. Access: API.
Unified ELO 1468 ± 35, rank #914 of 1776 rated models, from 42 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ConvRe - Re2Text Easy | 98.7 | Accuracy (%) | 100 |
| JustEval - Helpfulness | 4.9 | Score (1-5) | 100 |
| LLM Trustworthy - Out-of-Distribution | 87.55 | Trust Score (%) | 100 |
| HellaSwag | 95.3 | Accuracy (%) | 98 |
| SpeechMap Compliance | 95.2 | % Requests Completed | 97 |
| JustEval - Clarity | 4.99 | Score (1-5) | 96.7 |
| JustEval - Factuality | 4.9 | Score (1-5) | 96.7 |
| MMLU | 86.4 | Accuracy (%) | 95.6 |
| WinoGrande | 87.5 | Accuracy (%) | 95.6 |
| GSM8K | 92 | Accuracy (%) | 94.7 |
| LLM Trustworthy - Adversarial | 64.04 | Trust Score (%) | 92 |
| AlpacaEval 1.0 | 94.78 | Win Rate (%) | 91.6 |
Interactive version: theaggregate.ai/model?slug=gpt-4-0314 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.