Sci-Rho: leaderboard

Metric: Average-case accuracy (%) over all seven languages (English, Arabic, Chinese, Hindi, Indonesian, Kazakh, Swahili) on 606 expert-written executable templates per language (mathematics, physics, chemistry, biology and computer science, many drawn from olympiad problems), each rendered as 10 visually and numerically varied image-grounded instances; the mean of the per-language accuracies; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 17 models tracked.

Top models

#ModelScore
1GPT-5.4 (Non-reasoning)88.8
2Gemini 2.5 Pro85.8
3Qwen 3.5 27B (Non-reasoning)82.3
4Qwen 3.5 122B A10B (Non-reasoning)80.8
5Gemini 2.5 Flash (Non-reasoning)80.4
6Llama 4 Scout Instruct69.2
7Qwen 3 VL 8B Instruct65.3
8InternVL3.5-8B59.8
9Gemma 3 12B (IT)53.6
10Molmo2-8B43.9
11Gemma 3 4B (IT)36.4
12Qwen 3 VL 4B Instruct26.8

Interactive version: theaggregate.ai/benchmark?slug=sci-rho · How It Works · Data refreshed daily, snapshot 2026-09-29.