code-davinci-002: benchmark results

Codex-era OpenAI completion model trained for code generation and related tasks. Provider: OpenAI. Released 2022-03-15. Access: API.

Unified ELO 1424 ± 1, rank #1056 of 1392 rated models, from 10 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Classic - Synthetic Reasoning Abstract53.96Exact Match (%)100
HELM Classic - Dyck80.53Exact Match (%)98.5
HELM Classic - GSM8K56.77Exact Match (%)98.5
HELM Classic - Synthetic Reasoning Natural68.36F1 (%)98.5
HELM Classic - MATH41.02Equivalent (%)97.1
HELM Classic - bAbI68.62Exact Match (%)97.1
HELM Classic - MATH Chain-of-Thought43.33Equivalent (%)95.6
GSM8K56.8Accuracy (%)60.2
HELM Classic - LSAT0Exact Match (%)0.7
HELM Classic - LegalSupport0Exact Match (%)0.7

Interactive version: theaggregate.ai/model?slug=code-davinci-002 · How It Works · Data refreshed daily, snapshot 2026-09-05.