ScienceArena (Vision Routes): leaderboard
Metric: Weighted average (0-100): mean of the 13 contest scores, each normalized to 0-100 by its full score, over IBO 2023, APhO, EuPhO, USNCO, IPhO, IChO, CPhO and CChO 2025 and INChO, NBPhO, USAPhO, IPhO and IChO 2026; interleaved solving, one subpart at a time with earlier turns in context; open-ended physics and chemistry turns graded against the official solution and process-credit rubric by the authors' medalist-calibrated LLM-judge pipeline; the nine other contests use text-only question renderings for every model and the four 2026 contests (NBPhO, USAPhO, IPhO and IChO 2026) are solved from the original question-page images (vision-capable routes). Source: arxiv.org. Saturation forecast: Around December 2026. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 88.7 |
| 2 | Gemini 3.5 Flash | 86.1 |
| 3 | Gemini 3 Flash | 83.6 |
| 4 | GPT-5.5 | 78.8 |
| 5 | Claude Opus 4.7 | 75.8 |
| 6 | Qwen 3.6 Plus | 75.8 |
| 7 | Seed 2.0 Pro | 69.4 |
| 8 | GPT-5.4 | 67.3 |
Interactive version: theaggregate.ai/benchmark?slug=sciencearena-vision-routes · How It Works · Data refreshed daily, snapshot 2026-09-29.