DocOps (Claude Code): leaderboard

Metric: Task pass rate (%; 210 native-format document tasks in Excel, Word, PowerPoint and PDF (50 atomic edits, 40 compositional, 60 single-document and 60 cross-document workflows), passed only when the submitted artifact satisfies every structural, semantic and preservation predicate of a deterministic in-container verifier; run through Harbor; Claude Code harness without document skills; higher is better). Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.655.2
2DeepSeek V4 Pro43.3
3Qwen 3.6 27B38.6
4Qwen 3.6 35B A3B32.9
5Qwen 3.5 27B30
6Qwen 3.5 122B A10B29.5
7Qwen 3.5 35B A3B26.2
8Gemma 4 31B23.3
9Qwen 3.5 9B18.1
10GLM 4.5 Air12.9

Interactive version: theaggregate.ai/benchmark?slug=docops-claude-code · How It Works · Data refreshed daily, snapshot 2026-09-29.