DetailVerifyBench - Synthetic Hallucinations: leaderboard

Metric: Token-level F1 (%, times 100) of hallucination localization on the synthetic version (hallucinations injected into the corrected captions by Gemini 3 Flash and kept only when a text-only GPT-5.2 detector missed them, two adversarial rounds): the model reproduces a long caption (over 200 words on average) of one of 1,000 images in five domains (GUI, nature, chart, movie, poster) and wraps every hallucinated token in tags, scored per token against human annotations; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 14 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)66
2Gemini 3 Pro (Preview)64
3Qwen 3.5 397B A17B61
4Kimi K2.558
5Qwen 3.5 9B58
6Seed 2.0 Pro57
7Qwen 3.5 35B A3B52
8GPT-5.445
9GPT-5.243
10Qwen 3 VL 8B (Thinking)39
11Step3 VL 10B35
12Claude Opus 4.69
13MiMo-V2-Pro3

Interactive version: theaggregate.ai/benchmark?slug=detailverifybench-synthetic-hallucinations · How It Works · Data refreshed daily, snapshot 2026-10-07.