Roboflow Vision Evals - Visual Understanding: leaderboard

Roboflow Vision Evals benchmark for visual QA tasks such as reading text from photos, counting objects, spotting defects, and understanding documents.

Metric: Pass Rate (%). Source: playground.roboflow.com. Status: years away from saturation. 77 models tracked.

Top models

#ModelScore
1Qwen 3.5 35B A3B79.1
2Claude Fable 579.1
3Gemini 3.5 Flash79.1
4GPT-5.577.61
5GPT-5.477.61
6GPT-5.4 Mini77.61
7Qwen 3.5 122B A10B76.12
8GPT-5.6 Sol76.12
9GPT-5.6 Terra76.12
10Gemini 3.1 Pro (Preview)75.76
11Gemini 3 Flash74.63
12Qwen 3.5 Plus74.63
13GPT-5 Mini73.13
14Qwen 3.5 9B71.64
15Qwen 3.5 27B71.64

Interactive version: theaggregate.ai/benchmark?slug=roboflow-vision-evals-visual-understanding · How It Works · Data refreshed daily, snapshot 2026-09-05.