davinci — benchmark results
Original GPT-3 base Completions API model. Provider: OpenAI. Released 2020-06-11. Access: API.
Unified ELO 1218 ± 15, rank #1668 of 1776 rated models, from 34 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Classic - BLiMP | 83.97 | Exact Match (%) | 100 |
| HELM Classic - OpenbookQA | 58.6 | Exact Match (%) | 88.7 |
| HELM Classic - Entity Data Imputation | 83.59 | Exact Match (%) | 87.9 |
| HELM Classic - Dyck | 66.8 | Exact Match (%) | 77.9 |
| HELM Classic - BBQ | 39 | Exact Match (%) | 70.7 |
| HELM Classic - HellaSwag | 77.5 | Exact Match (%) | 67.7 |
| HELM Classic - MMLU | 42.24 | Exact Match (%) | 66.7 |
| HELM Classic - NaturalQuestions Closed Book | 32.86 | F1 (%) | 65.2 |
| HELM Classic - WikiFact | 30.61 | Exact Match (%) | 65.2 |
| HELM Classic - NarrativeQA | 68.69 | F1 (%) | 64.6 |
| HELM Classic - XSUM | 12.63 | ROUGE-2 (%) | 63.4 |
| HELM Classic - Synthetic Reasoning Abstract | 23.58 | Exact Match (%) | 61.8 |
Interactive version: theaggregate.ai/model?slug=davinci · How the rankings work · Data refreshed daily, snapshot 2026-07-22.