VeriOCRBench - Diagnosis Accuracy (Assisted): leaderboard

Metric: Diagnosis accuracy (%; share of VeriOCRBench's 1,600 trap-injected OCR reasoning tasks (8 trap types over visual, contextual, factual and logical premises, images from 8 OCR datasets, human-verified) on which the model both rejects the unexecutable task and names the ground-truth trap mechanism; judged by GPT-4o at temperature 0; model temperature 0 where supported; assisted paradigm: a meta-instruction asks the model to verify that the task is completable before answering). Source: arxiv.org. Saturation forecast: Around August 2027. 15 models tracked.

Top models

#ModelScore
1Qwen 3.5 397B A17B80.5
2GPT-576.94
3Claude Sonnet 4.576.12
4Qwen 3.5 35B A3B74.19
5Qwen 3.5 122B A10B73.81
6GPT-4o67.38
7InternVL3-38B67.31
8Llama 4 Maverick60.88
9Llama 4 Scout54.62
10Gemini 3.1 Pro (Preview)7.25

Interactive version: theaggregate.ai/benchmark?slug=veriocrbench-diagnosis-accuracy-assisted · How It Works · Data refreshed daily, snapshot 2026-09-26.