text-davinci-003 — benchmark results

InstructGPT completion model snapshot from the text-davinci series. Provider: OpenAI. Released 2022-11-28. Access: API.

Unified ELO 1381 ± 24, rank #1306 of 1776 rated models, from 54 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Classic - HellaSwag82.2Exact Match (%)100
HELM Classic - NaturalQuestions Open Book77.02F1 (%)100
HELM Classic - OpenbookQA64.6Exact Match (%)100
HELM Classic - QuAC52.48F1 (%)100
HELM Classic - Synthetic Reasoning Natural73.35F1 (%)100
ToolBench - VirtualHome25.1Task Score100
ToolBench - Home Search97Task Score98.8
HELM Classic - CivilComments68.44Exact Match (%)98.5
HELM Classic - RAFT75.91Exact Match (%)98.5
ToolBench - Tabletop66.7Task Score97.7
ToolBench - The Cat API98Task Score97.7
ToolBench - Trip Booking89.2Task Score97.7

Interactive version: theaggregate.ai/model?slug=text-davinci-003 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.