text-davinci-002 — benchmark results

InstructGPT completion model snapshot from the text-davinci series. Provider: OpenAI. Released 2022-03-15. Access: API.

Unified ELO 1364 ± 18, rank #1362 of 1776 rated models, from 39 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Classic - BBQ88.07Exact Match (%)100
HELM Classic - TruthfulQA60.96Exact Match (%)98.5
HELM Classic - WikiFact39.25Exact Match (%)97
HELM Classic - HellaSwag81.5Exact Match (%)96.8
HELM Classic - OpenbookQA59.4Exact Match (%)96.8
HELM Classic - Synthetic Reasoning Natural62.29F1 (%)95.6
HELM Classic - BoolQ87.7Exact Match (%)95.5
HELM Classic - CivilComments66.84Exact Match (%)95.5
HELM Classic - Entity Data Imputation84.16Exact Match (%)95.5
HELM Classic - Entity Matching93.08Exact Match (%)95.5
HELM Classic - NaturalQuestions Open Book71.32F1 (%)95.4
HELM Classic - MATH32.79Equivalent (%)94.1

Interactive version: theaggregate.ai/model?slug=text-davinci-002 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.