O1 (Low): benchmark results

Provider: OpenAI. Released 2024-12-05. Access: API.

Unified ELO 1637 ± 1, rank #409 of 3078 rated models, from 14 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AbstentionBench - underspecified context - GSM8K - F197.73Abstention F1 (%)100
AbstentionBench - underspecified context - GSM8K - Recall95.63Abstention Recall (%)100
AbstentionBench - underspecified context - MMLU Math - F183.33Abstention F1 (%)94.7
AbstentionBench - underspecified context - MMLU Math - Recall71.43Abstention Recall (%)92.1
AbstentionBench - underspecified context - MMLU Math - Precision100Abstention Precision (%)89.5
AbstentionBench - underspecified context - GPQA-Diamond - Precision100Abstention Precision (%)80
AbstentionBench - underspecified context - MMLU History - Precision100Abstention Precision (%)72.5
AbstentionBench - underspecified context - GPQA-Diamond - F176.92Abstention F1 (%)65
AbstentionBench - underspecified context - GPQA-Diamond - Recall62.5Abstention Recall (%)65
AbstentionBench - underspecified context - MMLU History - F152.83Abstention F1 (%)65
AbstentionBench - underspecified context - GSM8K - Precision99.91Abstention Precision (%)63.6
AbstentionBench - underspecified context - MMLU History - Recall35.9Abstention Recall (%)62.5

Interactive version: theaggregate.ai/model?slug=o1-low · How It Works · Data refreshed daily, snapshot 2026-09-19.