VeriOCRBench - Over-Refusal Rate (Assisted): leaderboard

Metric: Over-refusal rate (%; share of VeriOCRBench's 200 trap-free control tasks the model wrongly rejects; assisted paradigm: a meta-instruction asks the model to verify that the task is completable before answering; judged by GPT-4o at temperature 0; lower is better). Source: arxiv.org. Saturation forecast: Around December 2026. 15 models tracked.

Top models

#ModelScore
1GPT-59.5
2Qwen 3.5 122B A10B9.5
3Claude Sonnet 4.511.5
4Qwen 3.5 35B A3B12.5
5Qwen 3.5 397B A17B14.5
6GPT-4o15.5
7Llama 4 Scout16
8Llama 4 Maverick18.5
9InternVL3-38B32
10Gemini 3.1 Pro (Preview)79.5

Interactive version: theaggregate.ai/benchmark?slug=veriocrbench-over-refusal-rate-assisted · How It Works · Data refreshed daily, snapshot 2026-09-26.