DocFormBench (Docx-Skill): leaderboard

Metric: Format accuracy (%) on DocFormBench: share of required formatting attributes of the 500 content-aware Text-to-Format instances (Chinese and English .docx documents of 12 types) that exactly match the target, averaged per object type (page, paragraph, table, image) and then over types; the Docx-Skill baseline (Anthropic's open-source Word editing skill, invoked through the nanobot agent framework; default settings, at most 50 inference steps); reasoning (thinking) disabled; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 7 models tracked.

Top models

#ModelScore
1GLM-4.6V (Non-reasoning)25.57
2Gemini 3 Flash (Minimal)23.55
3DeepSeek V3.2 (Non-reasoning)20.71
4GLM-4.6 (Non-reasoning)18.19
5Qwen 3 VL 235B A22B Instruct15.53
6Qwen 3 235B A22B 2507 Instruct12.36
7GPT-5 (Minimal)4.49

Interactive version: theaggregate.ai/benchmark?slug=docformbench-docx-skill · How It Works · Data refreshed daily, snapshot 2026-09-29.