OCRBench v2 — leaderboard

OCRBench v2 evaluates large multimodal models on bilingual visual text localization and reasoning tasks.

Metric: Average (self-reported). Source: benchmarklist.com. Status: saturation imminent. 29 models tracked.

Top models

#ModelScore
1Qwen 3.5 9B64.1
2Gemini 2.5 Pro62.2
3Qwen3 Omni 30B A3B Instruct61.3
4Claude Opus 4.659.8
5Ovis2-8B56
6GPT-555.5
7GPT-5.252.6
8Gemini 1.5 Pro51.6
9Molmo2-8B44.7
10Ministral 3 14B44
11Phi-4 Multimodal Instruct38.1

Interactive version: theaggregate.ai/benchmark?slug=ocrbench-v2 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.