PPTBench: leaderboard
Metric: Mean reconstruction score (0-100): per task, the product of the artifact-validity, semantic-correctness and rendering-quality gates (majority of three judge rounds) and the layout (max 30), text (max 40) and local-graphics (max 30) scores left after severity-weighted defect deductions; invalid artifacts and gate failures score 0; 500 frozen tasks from arXiv flow diagrams, one run each; Agentic Judge GPT-5.6 Luna High, three rounds). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 31 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Kimi K3 (OpenCode, high) | 67.8 |
| 2 | GPT-5.6 Sol (Codex, max) | 49.28 |
| 3 | Qwen 3.8 Max (Claude Code, xhigh) | 47.69 |
| 4 | GPT-5.6 Sol (Codex, xhigh) | 36.24 |
| 5 | Claude Opus 5 (Claude Code, xhigh) | 35.91 |
| 6 | Claude Opus 5 (Claude Code, max) | 33.54 |
| 7 | GPT-5.6 Terra (Codex, max) | 33.46 |
| 8 | Claude Opus 5 (Claude Code, high) | 33.12 |
| 9 | GLM-5.3-Flash (OpenCode, max) | 32.42 |
| 10 | GPT-5.6 Luna (Codex, max) | 31.02 |
Interactive version: theaggregate.ai/benchmark?slug=pptbench · How It Works · Data refreshed daily, snapshot 2026-09-26.