davinci-002: benchmark results
Replacement GPT-3 base Completions API model for fine-tuning and legacy completions. Provider: OpenAI. Released 2023-08-22. Access: API.
Unified ELO 1320 ± 38, rank #2387 of 2656 rated models, from 25 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SAD - Anti-Imitation | 15.39 | Score (%, higher is better) | 80 |
| SAD - Anti-Imitation (Situating Prompt) | 15.39 | Score (%, higher is better) | 80 |
| SAD - Predict Words | 15.47 | Self-prediction score (%, higher is better) | 36.4 |
| SAD - Predict Words (Situating Prompt) | 14.2 | Self-prediction score (%, higher is better) | 31.8 |
| SAD - Self-Recognition | 50.56 | Score (%, higher is better) | 30 |
| SAD - Self-Recognition (Situating Prompt) | 50.63 | Score (%, higher is better) | 30 |
| SAD - Introspection | 23.98 | Score (%, higher is better) | 25 |
| SAD - Introspection (Situating Prompt) | 23.42 | Score (%, higher is better) | 25 |
| InfiBench | 21.25 | Score (%) | 13.3 |
| SAD - Stages | 36.69 | Score (%, higher is better) | 10 |
| SAD - Facts (Situating Prompt) | 38.65 | Score (%, higher is better) | 5 |
| SAD - Influence | 42.19 | Score (%, higher is better) | 5 |
Interactive version: theaggregate.ai/model?slug=davinci-002 · How It Works · Data refreshed daily, snapshot 2026-09-19.