GPT-3.5 Turbo (0301): benchmark results
Provider: OpenAI. Released 2023-03-01. Access: API.
Unified ELO 1439 ± 1, rank #994 of 1392 rated models, from 46 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ConvRe - Re2Text Hard | 59 | Accuracy (%) | 100 |
| HELM Classic - MATH | 48.83 | Equivalent (%) | 100 |
| HELM Classic - RAFT | 76.82 | Exact Match (%) | 100 |
| HELM Classic - Entity Matching | 94.26 | Exact Match (%) | 98.5 |
| HELM Classic - LSAT | 25.22 | Exact Match (%) | 98.5 |
| HELM Classic - MATH Chain-of-Thought | 68.94 | Equivalent (%) | 98.5 |
| HELM Classic - MMLU | 58.98 | Exact Match (%) | 98.5 |
| HELM Classic - QuAC | 51.2 | F1 (%) | 98.5 |
| HELM Classic - GSM8K | 53.1 | Exact Match (%) | 97.1 |
| HELM Classic - LegalSupport | 62.78 | Exact Match (%) | 97.1 |
| HELM Classic - Synthetic Reasoning Natural | 63.12 | F1 (%) | 97.1 |
| HELM Classic - CivilComments | 67.41 | Exact Match (%) | 97 |
Interactive version: theaggregate.ai/model?slug=gpt-3-5-turbo-0301 · How It Works · Data refreshed daily, snapshot 2026-09-05.