text-davinci-002: benchmark results

InstructGPT completion model snapshot from the text-davinci series. Provider: OpenAI. Released 2022-03-15. Access: API.

Unified ELO 1442 ± 1, rank #980 of 1392 rated models, from 46 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Classic - BBQ88.07Exact Match (%)100
HELM Classic - TruthfulQA60.96Exact Match (%)98.5
HELM Classic - WikiFact39.25Exact Match (%)97
HELM90.5Mean win rate (self-reported)96.9
HELM Classic - HellaSwag81.5Exact Match (%)96.8
HELM Classic - OpenbookQA59.4Exact Match (%)96.8
HELM Classic - Synthetic Reasoning Natural62.29F1 (%)95.6
HELM Classic - BoolQ87.7Exact Match (%)95.5
HELM Classic - CivilComments66.84Exact Match (%)95.5
HELM Classic - Entity Data Imputation84.16Exact Match (%)95.5
HELM Classic - Entity Matching93.08Exact Match (%)95.5
HELM Classic - NaturalQuestions Open Book71.32F1 (%)95.4

Interactive version: theaggregate.ai/model?slug=text-davinci-002 · How It Works · Data refreshed daily, snapshot 2026-09-05.