GPT-3.5 Turbo (0613) — benchmark results

Provider: OpenAI. Released 2023-06-13. Access: API.

Unified ELO 1377 ± 25, rank #1320 of 1776 rated models, from 60 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Classic - CivilComments69.64Exact Match (%)100
HELM Classic - MATH Chain-of-Thought71.86Equivalent (%)100
HELM Classic - MATH45.28Equivalent (%)98.5
HELM Classic - Synthetic Reasoning Abstract50.94Exact Match (%)98.5
FastEval65.17Total Score96.9
HELM Classic - QuAC48.5F1 (%)96.9
HELM Classic - RAFT74.77Exact Match (%)95.5
HELM Classic - Synthetic Reasoning Natural58.56F1 (%)94.1
HELM Classic - GSM8K46.9Exact Match (%)92.6
HELM Classic - Entity Matching91.77Exact Match (%)92.4
HELM Classic - BoolQ87Exact Match (%)90.9
InfiBench56.47Score (%)88.6

Interactive version: theaggregate.ai/model?slug=gpt-3-5-turbo-0613 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.