O3 (Low): benchmark results

O3 evaluated at the low reasoning-effort setting. Provider: OpenAI. Released 2025-04-16. Access: API.

Unified ELO 1617 ± 1, rank #367 of 1761 rated models, from 14 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
UGI - Natural Intelligence65.53NatInt Score96.2
UGI - Writing65.79Writing Score96.1
UGI Leaderboard52.09UGI Score93.2
LLM Chess (Saplin)738.6ELO80.7
Chess Puzzles (Epoch AI)27Accuracy (%)76.3
DROP82.3F1 Score63.3
OTIS Mock AIME 2024-2560Accuracy (%)45.5
FrontierMath - Tiers 1-39.72Accuracy (%, 290 problems)44.4
ARC-AGI-141.5Accuracy (%)35.9
FrontierMath - Tiers 1-3 (v2)19.3Accuracy (%, 285 private v2 problems)31.5
ARC-AGI-21.99Accuracy (%)25.7
UGI - Willingness (W/10)2.2W/10 Score15.4

Interactive version: theaggregate.ai/model?slug=o3-low · How It Works · Data refreshed daily, snapshot 2026-09-05.