text-babbage-001 — benchmark results

Provider: OpenAI. Released 2022-01-27. Access: API.

Unified ELO 1102 ± 30, rank #1761 of 1776 rated models, from 34 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Classic - CNN/DailyMail15.11ROUGE-2 (%)75.6
HELM Classic - Synthetic Reasoning Natural21.16F1 (%)61.8
HELM Classic - MS MARCO TREC44.92NDCG@10 (%)56.7
HELM Classic - BBQ35.9Exact Match (%)52.4
HELM Classic - TruthfulQA23.29Exact Match (%)51.5
HELM Classic - LegalSupport51.74Exact Match (%)46.3
HELM Classic - Entity Matching73.36Exact Match (%)42.4
HELM Classic - MS MARCO Regular20.76RR@10 (%)41.4
HELM Classic - LSAT18.99Exact Match (%)33.1
HELM Classic - IMDB91.27Exact Match (%)28.8
HELM Classic - RAFT50.91Exact Match (%)22.7
HELM Classic - Entity Data Imputation70.4Exact Match (%)21.2

Interactive version: theaggregate.ai/model?slug=text-babbage-001 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.