ScienceArena (Vision Routes): leaderboard

Metric: Weighted average (0-100): mean of the 13 contest scores, each normalized to 0-100 by its full score, over IBO 2023, APhO, EuPhO, USNCO, IPhO, IChO, CPhO and CChO 2025 and INChO, NBPhO, USAPhO, IPhO and IChO 2026; interleaved solving, one subpart at a time with earlier turns in context; open-ended physics and chemistry turns graded against the official solution and process-credit rubric by the authors' medalist-calibrated LLM-judge pipeline; the nine other contests use text-only question renderings for every model and the four 2026 contests (NBPhO, USAPhO, IPhO and IChO 2026) are solved from the original question-page images (vision-capable routes). Source: arxiv.org. Saturation forecast: Around December 2026. 8 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)88.7
2Gemini 3.5 Flash86.1
3Gemini 3 Flash83.6
4GPT-5.578.8
5Claude Opus 4.775.8
6Qwen 3.6 Plus75.8
7Seed 2.0 Pro69.4
8GPT-5.467.3

Interactive version: theaggregate.ai/benchmark?slug=sciencearena-vision-routes · How It Works · Data refreshed daily, snapshot 2026-09-29.