text-davinci-003 — benchmark results
InstructGPT completion model snapshot from the text-davinci series. Provider: OpenAI. Released 2022-11-28. Access: API.
Unified ELO 1381 ± 24, rank #1306 of 1776 rated models, from 54 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Classic - HellaSwag | 82.2 | Exact Match (%) | 100 |
| HELM Classic - NaturalQuestions Open Book | 77.02 | F1 (%) | 100 |
| HELM Classic - OpenbookQA | 64.6 | Exact Match (%) | 100 |
| HELM Classic - QuAC | 52.48 | F1 (%) | 100 |
| HELM Classic - Synthetic Reasoning Natural | 73.35 | F1 (%) | 100 |
| ToolBench - VirtualHome | 25.1 | Task Score | 100 |
| ToolBench - Home Search | 97 | Task Score | 98.8 |
| HELM Classic - CivilComments | 68.44 | Exact Match (%) | 98.5 |
| HELM Classic - RAFT | 75.91 | Exact Match (%) | 98.5 |
| ToolBench - Tabletop | 66.7 | Task Score | 97.7 |
| ToolBench - The Cat API | 98 | Task Score | 97.7 |
| ToolBench - Trip Booking | 89.2 | Task Score | 97.7 |
Interactive version: theaggregate.ai/model?slug=text-davinci-003 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.