DetailVerifyBench: leaderboard
Metric: Token-level F1 (%, times 100) of hallucination localization on the real-hallucination version (Gemini 3 Pro captions corrected by human annotators; the hallucinated spans are the corrected ones): the model reproduces a long caption (over 200 words on average) of one of 1,000 images in five domains (GUI, nature, chart, movie, poster) and wraps every hallucinated token in tags, scored per token against human annotations; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 14 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3.5 397B A17B | 13 |
| 2 | Gemini 3.1 Pro (Preview) | 12 |
| 3 | Seed 2.0 Pro | 11 |
| 4 | Qwen 3.5 35B A3B | 10 |
| 5 | Kimi K2.5 | 8 |
| 6 | Qwen 3.5 9B | 8 |
| 7 | GPT-5.4 | 7 |
| 8 | Claude Opus 4.6 | 7 |
| 9 | GPT-5.2 | 5 |
| 10 | Qwen 3 VL 8B (Thinking) | 3 |
| 11 | Step3 VL 10B | 3 |
| 12 | MiMo-V2-Pro | 1 |
Interactive version: theaggregate.ai/benchmark?slug=detailverifybench · How It Works · Data refreshed daily, snapshot 2026-10-07.