davinci — benchmark results

Original GPT-3 base Completions API model. Provider: OpenAI. Released 2020-06-11. Access: API.

Unified ELO 1218 ± 15, rank #1668 of 1776 rated models, from 34 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Classic - BLiMP83.97Exact Match (%)100
HELM Classic - OpenbookQA58.6Exact Match (%)88.7
HELM Classic - Entity Data Imputation83.59Exact Match (%)87.9
HELM Classic - Dyck66.8Exact Match (%)77.9
HELM Classic - BBQ39Exact Match (%)70.7
HELM Classic - HellaSwag77.5Exact Match (%)67.7
HELM Classic - MMLU42.24Exact Match (%)66.7
HELM Classic - NaturalQuestions Closed Book32.86F1 (%)65.2
HELM Classic - WikiFact30.61Exact Match (%)65.2
HELM Classic - NarrativeQA68.69F1 (%)64.6
HELM Classic - XSUM12.63ROUGE-2 (%)63.4
HELM Classic - Synthetic Reasoning Abstract23.58Exact Match (%)61.8

Interactive version: theaggregate.ai/model?slug=davinci · How the rankings work · Data refreshed daily, snapshot 2026-07-22.