ScienceArena - APhO 2025: leaderboard
Metric: Score (out of 30): points awarded on the Asian Physics Olympiad 2025 theory problems, text-only question renderings; interleaved solving, one subpart at a time with earlier turns in context; open-ended physics and chemistry turns graded against the official solution and process-credit rubric by the authors' medalist-calibrated LLM-judge pipeline. Source: arxiv.org. Saturation forecast: Estimated already saturated. 14 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 28.5 |
| 2 | GLM-5.1 | 28.4 |
| 3 | Qwen 3.7 Max | 28.1 |
| 4 | Gemini 3.5 Flash | 27.91 |
| 5 | Qwen 3.6 Plus | 27.56 |
| 6 | Gemini 3 Flash | 26.75 |
| 7 | DeepSeek V4 Pro | 26.01 |
| 8 | Claude Opus 4.7 | 24.19 |
| 9 | MiniMax-M2.7 | 22.56 |
| 10 | GPT-5.4 | 22.01 |
| 11 | Seed 2.0 Pro | 21.72 |
| 12 | GPT-5.5 | 21.5 |
| 13 | MiMo-V2.5-Pro | 21.33 |
| 14 | DeepSeek V4 Flash | 20.44 |
Interactive version: theaggregate.ai/benchmark?slug=sciencearena-apho-2025 · How It Works · Data refreshed daily, snapshot 2026-09-29.