code-davinci-002 — benchmark results
Codex-era OpenAI completion model trained for code generation and related tasks. Provider: OpenAI. Released 2022-03-15. Access: API.
Unified ELO 1206 ± 164, rank #1685 of 1776 rated models, from 10 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Classic - Synthetic Reasoning Abstract | 53.96 | Exact Match (%) | 100 |
| HELM Classic - Dyck | 80.53 | Exact Match (%) | 98.5 |
| HELM Classic - GSM8K | 56.77 | Exact Match (%) | 98.5 |
| HELM Classic - Synthetic Reasoning Natural | 68.36 | F1 (%) | 98.5 |
| HELM Classic - MATH | 41.02 | Equivalent (%) | 97.1 |
| HELM Classic - bAbI | 68.62 | Exact Match (%) | 97.1 |
| HELM Classic - MATH Chain-of-Thought | 43.33 | Equivalent (%) | 95.6 |
| GSM8K | 56.8 | Accuracy (%) | 59.6 |
| HELM Classic - LSAT | 0 | Exact Match (%) | 0.7 |
| HELM Classic - LegalSupport | 0 | Exact Match (%) | 0.7 |
Interactive version: theaggregate.ai/model?slug=code-davinci-002 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.