DetailVerifyBench: leaderboard

Metric: Token-level F1 (%, times 100) of hallucination localization on the real-hallucination version (Gemini 3 Pro captions corrected by human annotators; the hallucinated spans are the corrected ones): the model reproduces a long caption (over 200 words on average) of one of 1,000 images in five domains (GUI, nature, chart, movie, poster) and wraps every hallucinated token in tags, scored per token against human annotations; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 14 models tracked.

Top models

#ModelScore
1Qwen 3.5 397B A17B13
2Gemini 3.1 Pro (Preview)12
3Seed 2.0 Pro11
4Qwen 3.5 35B A3B10
5Kimi K2.58
6Qwen 3.5 9B8
7GPT-5.47
8Claude Opus 4.67
9GPT-5.25
10Qwen 3 VL 8B (Thinking)3
11Step3 VL 10B3
12MiMo-V2-Pro1

Interactive version: theaggregate.ai/benchmark?slug=detailverifybench · How It Works · Data refreshed daily, snapshot 2026-10-07.