Sci-Rho - Mathematics (English): leaderboard

Metric: Average-case accuracy (%) on the English mathematics templates of Sci-Rho: 606 expert-written executable templates per language (mathematics, physics, chemistry, biology and computer science, many drawn from olympiad problems), each rendered as 10 visually and numerically varied image-grounded instances; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 17 models tracked.

Top models

#ModelScore
1GPT-5.4 (Non-reasoning)90.3
2Gemini 2.5 Pro86.2
3Gemini 2.5 Flash (Non-reasoning)85.2
4Qwen 3.5 27B (Non-reasoning)85.1
5Qwen 3.5 122B A10B (Non-reasoning)82.1
6Llama 4 Scout Instruct71.4
7InternVL3.5-8B64.8
8Gemma 3 12B (IT)61.7
9Qwen 3 VL 8B Instruct61.5
10Molmo2-8B46.5
11Gemma 3 4B (IT)46.2
12Qwen 3 VL 4B Instruct31.7

Interactive version: theaggregate.ai/benchmark?slug=sci-rho-mathematics-english · How It Works · Data refreshed daily, snapshot 2026-09-29.