ScienceArena - APhO 2025: leaderboard

Metric: Score (out of 30): points awarded on the Asian Physics Olympiad 2025 theory problems, text-only question renderings; interleaved solving, one subpart at a time with earlier turns in context; open-ended physics and chemistry turns graded against the official solution and process-credit rubric by the authors' medalist-calibrated LLM-judge pipeline. Source: arxiv.org. Saturation forecast: Estimated already saturated. 14 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)28.5
2GLM-5.128.4
3Qwen 3.7 Max28.1
4Gemini 3.5 Flash27.91
5Qwen 3.6 Plus27.56
6Gemini 3 Flash26.75
7DeepSeek V4 Pro26.01
8Claude Opus 4.724.19
9MiniMax-M2.722.56
10GPT-5.422.01
11Seed 2.0 Pro21.72
12GPT-5.521.5
13MiMo-V2.5-Pro21.33
14DeepSeek V4 Flash20.44

Interactive version: theaggregate.ai/benchmark?slug=sciencearena-apho-2025 · How It Works · Data refreshed daily, snapshot 2026-09-29.