DocFormBench (Docx-Skill): leaderboard
Metric: Format accuracy (%) on DocFormBench: share of required formatting attributes of the 500 content-aware Text-to-Format instances (Chinese and English .docx documents of 12 types) that exactly match the target, averaged per object type (page, paragraph, table, image) and then over types; the Docx-Skill baseline (Anthropic's open-source Word editing skill, invoked through the nanobot agent framework; default settings, at most 50 inference steps); reasoning (thinking) disabled; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GLM-4.6V (Non-reasoning) | 25.57 |
| 2 | Gemini 3 Flash (Minimal) | 23.55 |
| 3 | DeepSeek V3.2 (Non-reasoning) | 20.71 |
| 4 | GLM-4.6 (Non-reasoning) | 18.19 |
| 5 | Qwen 3 VL 235B A22B Instruct | 15.53 |
| 6 | Qwen 3 235B A22B 2507 Instruct | 12.36 |
| 7 | GPT-5 (Minimal) | 4.49 |
Interactive version: theaggregate.ai/benchmark?slug=docformbench-docx-skill · How It Works · Data refreshed daily, snapshot 2026-09-29.