Sci-Rho - Worst-Case: leaderboard

Metric: Worst-case accuracy (%) over all seven languages on 606 expert-written executable templates per language (mathematics, physics, chemistry, biology and computer science, many drawn from olympiad problems), each rendered as 10 visually and numerically varied image-grounded instances: the share of templates a model answers correctly on every one of its 10 generated variants, averaged over languages; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 17 models tracked.

Top models

#ModelScore
1GPT-5.4 (Non-reasoning)73.9
2Gemini 2.5 Pro62.5
3Gemini 2.5 Flash (Non-reasoning)56.6
4Qwen 3.5 27B (Non-reasoning)51.2
5Qwen 3.5 122B A10B (Non-reasoning)43.6
6Llama 4 Scout Instruct33.7
7Qwen 3 VL 8B Instruct29.5
8InternVL3.5-8B25.8
9Gemma 3 12B (IT)18.7
10Molmo2-8B15.5
11Gemma 3 4B (IT)7.2
12Qwen 3 VL 4B Instruct3.3

Interactive version: theaggregate.ai/benchmark?slug=sci-rho-worst-case · How It Works · Data refreshed daily, snapshot 2026-09-29.