text-curie-001 — benchmark results
Provider: OpenAI. Released 2022-01-27. Access: API.
Unified ELO 1137 ± 28, rank #1748 of 1776 rated models, from 44 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Classic - CNN/DailyMail | 15.16 | ROUGE-2 (%) | 79.3 |
| HELM Classic - Entity Matching | 85.2 | Exact Match (%) | 75.8 |
| HELM Classic - MS MARCO TREC | 50.72 | NDCG@10 (%) | 73.3 |
| HELM Classic - Synthetic Reasoning Natural | 22.06 | F1 (%) | 66.2 |
| HELM Classic - TruthfulQA | 25.74 | Exact Match (%) | 63.6 |
| HELM Classic - MS MARCO Regular | 27.12 | RR@10 (%) | 62.1 |
| HELM Classic - Entity Data Imputation | 79.13 | Exact Match (%) | 60.6 |
| HELM Classic - QuAC | 35.79 | F1 (%) | 52.3 |
| ToolBench - WebShop Long | 0 | Task Score | 44.2 |
| HELM Classic - CivilComments | 53.73 | Exact Match (%) | 43.9 |
| HELM Classic - Synthetic Reasoning Abstract | 18.99 | Exact Match (%) | 33.8 |
| HELM Classic - IMDB | 92.27 | Exact Match (%) | 33.3 |
Interactive version: theaggregate.ai/model?slug=text-curie-001 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.