O3 (Low) — benchmark results

O3 evaluated at the low reasoning-effort setting. Provider: OpenAI. Released 2025-04-16. Access: API.

Unified ELO 1708 ± 39, rank #261 of 1776 rated models, from 11 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
UGI - Natural Intelligence65.53NatInt Score96.9
UGI - Writing65.79Writing Score96.7
UGI Leaderboard52.09UGI Score93.3
LLM Chess (Saplin)738.6ELO87.9
Epoch AI - ECI147.07ECI Score66
DROP82.3F1 Score63.3
FrontierMath - Tiers 1-39.72Accuracy (%, 290 problems)44.4
ARC-AGI-141.5Accuracy (%)43.5
ARC-AGI-21.99Accuracy (%)31.9
UGI - Willingness (W/10)2.2W/10 Score15.2
GDPval (OpenAI Evals)11.8Win Rate (%)6.7

Interactive version: theaggregate.ai/model?slug=o3-low · How the rankings work · Data refreshed daily, snapshot 2026-07-22.