DetailVerifyBench - Synthetic Hallucinations: leaderboard
Metric: Token-level F1 (%, times 100) of hallucination localization on the synthetic version (hallucinations injected into the corrected captions by Gemini 3 Flash and kept only when a text-only GPT-5.2 detector missed them, two adversarial rounds): the model reproduces a long caption (over 200 words on average) of one of 1,000 images in five domains (GUI, nature, chart, movie, poster) and wraps every hallucinated token in tags, scored per token against human annotations; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 14 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 66 |
| 2 | Gemini 3 Pro (Preview) | 64 |
| 3 | Qwen 3.5 397B A17B | 61 |
| 4 | Kimi K2.5 | 58 |
| 5 | Qwen 3.5 9B | 58 |
| 6 | Seed 2.0 Pro | 57 |
| 7 | Qwen 3.5 35B A3B | 52 |
| 8 | GPT-5.4 | 45 |
| 9 | GPT-5.2 | 43 |
| 10 | Qwen 3 VL 8B (Thinking) | 39 |
| 11 | Step3 VL 10B | 35 |
| 12 | Claude Opus 4.6 | 9 |
| 13 | MiMo-V2-Pro | 3 |
Interactive version: theaggregate.ai/benchmark?slug=detailverifybench-synthetic-hallucinations · How It Works · Data refreshed daily, snapshot 2026-10-07.