DocOps (Claude Code, Skills): leaderboard

Metric: Task pass rate (%; 210 native-format document tasks in Excel, Word, PowerPoint and PDF (50 atomic edits, 40 compositional, 60 single-document and 60 cross-document workflows), passed only when the submitted artifact satisfies every structural, semantic and preservation predicate of a deterministic in-container verifier; run through Harbor; Claude Code harness with the Anthropic xlsx, docx, pptx and pdf skills; higher is better). Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.651.9
2Qwen 3.6 27B42.4
3DeepSeek V4 Pro40
4Qwen 3.5 27B37.1
5Qwen 3.6 35B A3B33.8
6Gemma 4 31B30
7Qwen 3.5 35B A3B29.5
8Qwen 3.5 122B A10B29
9Qwen 3.5 9B17.6
10GLM 4.5 Air16.7

Interactive version: theaggregate.ai/benchmark?slug=docops-claude-code-skills · How It Works · Data refreshed daily, snapshot 2026-09-29.