GPT-4 Turbo (Preview) — benchmark results
OpenAI preview GPT-4 Turbo model row, tracked separately from the GA GPT-4 Turbo row. Provider: OpenAI. Released 2023-11-06. Access: API.
Unified ELO 1595 ± 32, rank #486 of 1776 rated models, from 21 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| T-Eval | 86.4 | Overall Score (%) | 100 |
| InfiBench | 68.42 | Score (%) | 99 |
| Klu LLM Leaderboard | 100 | Klu Index | 98.8 |
| EvalPlus (HumanEval+ & MBPP+) | 77.5 | Pass@1 avg (%) | 91.1 |
| LLM-AggreFact | 76.16 | Balanced Accuracy (%) | 89.5 |
| HELM NaturalQuestions (Open) | 76.29 | F1 (%) | 82.2 |
| HELM NaturalQuestions (Closed) | 43.47 | F1 (%) | 77.8 |
| SEAL - Korean | 60.76 | Score | 77.8 |
| HELM (Stanford) | 69.82 | Mean Win Rate (%) | 74.4 |
| SEAL - Math | 95.1 | Score | 73.3 |
| SEAL - Adversarial Robustness | 20 | Score | 71.4 |
| SpeechMap Compliance | 75.1 | % Requests Completed | 68.5 |
Interactive version: theaggregate.ai/model?slug=gpt-4-turbo-preview · How the rankings work · Data refreshed daily, snapshot 2026-07-22.