GPT-3.5 Turbo (0301) — benchmark results
Provider: OpenAI. Released 2023-03-01. Access: API.
Unified ELO 1348 ± 24, rank #1423 of 1776 rated models, from 44 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ConvRe - Re2Text Hard | 59 | Accuracy (%) | 100 |
| HELM Classic - MATH | 48.83 | Equivalent (%) | 100 |
| HELM Classic - RAFT | 76.82 | Exact Match (%) | 100 |
| HELM Classic - Entity Matching | 94.26 | Exact Match (%) | 98.5 |
| HELM Classic - LSAT | 25.22 | Exact Match (%) | 98.5 |
| HELM Classic - MATH Chain-of-Thought | 68.94 | Equivalent (%) | 98.5 |
| HELM Classic - MMLU | 58.98 | Exact Match (%) | 98.5 |
| HELM Classic - QuAC | 51.2 | F1 (%) | 98.5 |
| HELM Classic - GSM8K | 53.1 | Exact Match (%) | 97.1 |
| HELM Classic - LegalSupport | 62.78 | Exact Match (%) | 97.1 |
| HELM Classic - Synthetic Reasoning Natural | 63.12 | F1 (%) | 97.1 |
| HELM Classic - CivilComments | 67.41 | Exact Match (%) | 97 |
Interactive version: theaggregate.ai/model?slug=gpt-3-5-turbo-0301 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.