VeriOCRBench - Visual Dimension (Proactive): leaderboard

Metric: Diagnosis accuracy (%; the 400 trap-injected tasks of the visual verification dimension (image degradation and occluded target: the required text or visual evidence is unreadable or hidden); the model must reject the task and name the ground-truth trap mechanism; judged by GPT-4o at temperature 0; proactive paradigm: no verification instruction). Source: arxiv.org. Saturation forecast: Around 2029. 15 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.528.75
2Gemini 3.1 Pro (Preview)23.25
3GPT-520.5
4Qwen 3.5 35B A3B18.5
5Llama 4 Maverick14.75
6Qwen 3.5 397B A17B14.5
7Llama 4 Scout11.5
8GPT-4o11.25
9Qwen 3.5 122B A10B10.5
10InternVL3-38B6.75

Interactive version: theaggregate.ai/benchmark?slug=veriocrbench-visual-dimension-proactive · How It Works · Data refreshed daily, snapshot 2026-09-26.