VAMPS - Persian: leaderboard

Metric: Accuracy (%) on the original Persian versions of the 218 Konkour-seed questions (Iranian university entrance exam algebra and calculus, chosen so that plotting reveals the answer; four-option multiple choice), tool-enabled visual solving: the model must obtain one to four Desmos plots and answer only from the visible graph evidence, no analytical derivation; thinking disabled for every model, temperature 0, answers read from a required JSON block; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 18 models tracked.

Top models

#ModelScore
1Claude Opus 4.792.7
2Claude Sonnet 4.688.5
3Qwen 3.5 397B A17B (Non-reasoning)88.5
4Gemini 2.5 Flash (Non-reasoning)84.4
5Gemma 4 31B (Non-reasoning)84.4
6Qwen 3.5 27B (Non-reasoning)82.6
7Gemma 4 26B A4B (Non-reasoning)81.2
8Qwen 3.5 35B A3B (Non-reasoning)76.6
9Qwen 3 VL 32B Instruct71.6
10GPT-5.4 (Non-reasoning)70.6
11Ministral 3 14B53.2
12Ministral 3 8B50.5
13GPT-4o49.5
14Qwen 3 VL 8B Instruct49.1
15Gemma 3 27B46.8

Interactive version: theaggregate.ai/benchmark?slug=vamps-persian · How It Works · Data refreshed daily, snapshot 2026-09-29.