VeriOCRBench - Diagnosis Accuracy (Proactive): leaderboard

Metric: Diagnosis accuracy (%; share of VeriOCRBench's 1,600 trap-injected OCR reasoning tasks (8 trap types over visual, contextual, factual and logical premises, images from 8 OCR datasets, human-verified) on which the model both rejects the unexecutable task and names the ground-truth trap mechanism; judged by GPT-4o at temperature 0; model temperature 0 where supported; proactive paradigm: the standard image, textual premise and question with no verification instruction). Source: arxiv.org. Saturation forecast: Around 2030. 15 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.546.88
2Qwen 3.5 35B A3B25.81
3Gemini 3.1 Pro (Preview)24.69
4Qwen 3.5 122B A10B23.56
5Qwen 3.5 397B A17B23.44
6Llama 4 Scout21.06
7Llama 4 Maverick17.44
8GPT-4o13.19
9InternVL3-38B12.69
10GPT-511.5

Interactive version: theaggregate.ai/benchmark?slug=veriocrbench-diagnosis-accuracy-proactive · How It Works · Data refreshed daily, snapshot 2026-09-26.