DocFormBench (DocFormFlow) - Hallucinated Formatting: leaderboard
Metric: Hallucinated formatting rate (%) on DocFormBench: share of applied formatting changes that fall outside the intended scope (1 - precision), DocFormFlow workflow (localize targets, then call formatting tools; temperature 1.0, up to 3 verification and self-repair rounds; multimodal models also receive page images); reasoning disabled; lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GLM-4.6 (Non-reasoning) | 1.24 |
| 2 | DeepSeek V4 Pro (Non-reasoning) | 1.57 |
| 3 | Gemini 3 Flash (Minimal) | 1.58 |
| 4 | DeepSeek V4 Flash (Non-reasoning) | 1.8 |
| 5 | Qwen 3 VL 32B Instruct | 1.94 |
| 6 | DeepSeek V3.2 (Non-reasoning) | 1.99 |
| 7 | Qwen 3 VL 235B A22B Instruct | 2.03 |
| 8 | GLM-4.6V (Non-reasoning) | 2.14 |
| 9 | Qwen 3 VL 8B Instruct | 2.45 |
| 10 | Qwen 3 235B A22B 2507 Instruct | 2.69 |
| 11 | Qwen 3.5 397B A17B (Non-reasoning) | 2.69 |
| 12 | Qwen 3.5 122B A10B (Non-reasoning) | 3.06 |
| 13 | GPT-5 (Minimal) | 3.44 |
| 14 | Qwen 3.5 35B A3B (Non-reasoning) | 3.45 |
| 15 | Qwen 3 VL 30B A3B Instruct | 3.8 |
Interactive version: theaggregate.ai/benchmark?slug=docformbench-docformflow-hallucinated-formatting · How It Works · Data refreshed daily, snapshot 2026-09-29.