code-davinci-002 — benchmark results

Codex-era OpenAI completion model trained for code generation and related tasks. Provider: OpenAI. Released 2022-03-15. Access: API.

Unified ELO 1206 ± 164, rank #1685 of 1776 rated models, from 10 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Classic - Synthetic Reasoning Abstract53.96Exact Match (%)100
HELM Classic - Dyck80.53Exact Match (%)98.5
HELM Classic - GSM8K56.77Exact Match (%)98.5
HELM Classic - Synthetic Reasoning Natural68.36F1 (%)98.5
HELM Classic - MATH41.02Equivalent (%)97.1
HELM Classic - bAbI68.62Exact Match (%)97.1
HELM Classic - MATH Chain-of-Thought43.33Equivalent (%)95.6
GSM8K56.8Accuracy (%)59.6
HELM Classic - LSAT0Exact Match (%)0.7
HELM Classic - LegalSupport0Exact Match (%)0.7

Interactive version: theaggregate.ai/model?slug=code-davinci-002 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.