davinci-002: benchmark results

Replacement GPT-3 base Completions API model for fine-tuning and legacy completions. Provider: OpenAI. Released 2023-08-22. Access: API.

Unified ELO 1320 ± 38, rank #2387 of 2656 rated models, from 25 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
SAD - Anti-Imitation15.39Score (%, higher is better)80
SAD - Anti-Imitation (Situating Prompt)15.39Score (%, higher is better)80
SAD - Predict Words15.47Self-prediction score (%, higher is better)36.4
SAD - Predict Words (Situating Prompt)14.2Self-prediction score (%, higher is better)31.8
SAD - Self-Recognition50.56Score (%, higher is better)30
SAD - Self-Recognition (Situating Prompt)50.63Score (%, higher is better)30
SAD - Introspection23.98Score (%, higher is better)25
SAD - Introspection (Situating Prompt)23.42Score (%, higher is better)25
InfiBench21.25Score (%)13.3
SAD - Stages36.69Score (%, higher is better)10
SAD - Facts (Situating Prompt)38.65Score (%, higher is better)5
SAD - Influence42.19Score (%, higher is better)5

Interactive version: theaggregate.ai/model?slug=davinci-002 · How It Works · Data refreshed daily, snapshot 2026-09-19.