GPT-3.5 Turbo (0301) — benchmark results

Provider: OpenAI. Released 2023-03-01. Access: API.

Unified ELO 1348 ± 24, rank #1423 of 1776 rated models, from 44 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
ConvRe - Re2Text Hard59Accuracy (%)100
HELM Classic - MATH48.83Equivalent (%)100
HELM Classic - RAFT76.82Exact Match (%)100
HELM Classic - Entity Matching94.26Exact Match (%)98.5
HELM Classic - LSAT25.22Exact Match (%)98.5
HELM Classic - MATH Chain-of-Thought68.94Equivalent (%)98.5
HELM Classic - MMLU58.98Exact Match (%)98.5
HELM Classic - QuAC51.2F1 (%)98.5
HELM Classic - GSM8K53.1Exact Match (%)97.1
HELM Classic - LegalSupport62.78Exact Match (%)97.1
HELM Classic - Synthetic Reasoning Natural63.12F1 (%)97.1
HELM Classic - CivilComments67.41Exact Match (%)97

Interactive version: theaggregate.ai/model?slug=gpt-3-5-turbo-0301 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.