O3 (2025-04-16) (High) — benchmark results

O3 (2025-04-16) evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2025-04-16. Access: API.

Unified ELO 1754 ± 29, rank #190 of 1776 rated models, from 20 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
UGI - Writing66.75Writing Score97.1
UGI - Natural Intelligence66.19NatInt Score97
MATH Level 597.77Accuracy (%)95.4
SEAL - EnigmaEval11.91Score87.5
UGI Leaderboard46.52UGI Score83.7
SimpleQA Verified53Accuracy (%)76.6
MultiNRC45.5Score (self-reported)71.1
SEAL - MultiNRC45.5Score69.8
OTIS Mock AIME 2024-2583.89Accuracy (%)67.3
LiveCodeBench Pro1010Rating (CF-style)52.8
OpenCompass Language - Dialogue96.1Score (%)52
SEAL - TutorBench52.09Score50

Interactive version: theaggregate.ai/model?slug=o3-2025-04-16-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.