DocFormBench (DocFormFlow): leaderboard

Metric: Format accuracy (%) on DocFormBench: share of required formatting attributes of the 500 content-aware Text-to-Format instances (Chinese and English .docx documents of 12 types) that exactly match the target, averaged per object type (page, paragraph, table, image) and then over types; the authors' DocFormFlow workflow (localize targets, then call formatting tools; temperature 1.0, up to 3 verification and self-repair rounds; multimodal models also receive page images, language models the text only); reasoning (thinking) disabled; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 15 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Minimal)80.36
2GLM-4.6 (Non-reasoning)75.72
3Qwen 3.5 397B A17B (Non-reasoning)75.19
4DeepSeek V4 Pro (Non-reasoning)74
5Qwen 3 VL 32B Instruct73.04
6DeepSeek V3.2 (Non-reasoning)72.87
7GPT-5 (Minimal)72.53
8Qwen 3 235B A22B 2507 Instruct71.18
9Qwen 3 VL 235B A22B Instruct70.94
10Qwen 3.5 122B A10B (Non-reasoning)70.87
11GLM-4.6V (Non-reasoning)65.78
12Qwen 3 VL 30B A3B Instruct64.47
13DeepSeek V4 Flash (Non-reasoning)64.36
14Qwen 3.5 35B A3B (Non-reasoning)62.1
15Qwen 3 VL 8B Instruct47.64

Interactive version: theaggregate.ai/benchmark?slug=docformbench-docformflow · How It Works · Data refreshed daily, snapshot 2026-09-29.