PPTBench: leaderboard

Metric: Mean reconstruction score (0-100): per task, the product of the artifact-validity, semantic-correctness and rendering-quality gates (majority of three judge rounds) and the layout (max 30), text (max 40) and local-graphics (max 30) scores left after severity-weighted defect deductions; invalid artifacts and gate failures score 0; 500 frozen tasks from arXiv flow diagrams, one run each; Agentic Judge GPT-5.6 Luna High, three rounds). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 31 models tracked.

Top models

#ModelScore
1Kimi K3 (OpenCode, high)67.8
2GPT-5.6 Sol (Codex, max)49.28
3Qwen 3.8 Max (Claude Code, xhigh)47.69
4GPT-5.6 Sol (Codex, xhigh)36.24
5Claude Opus 5 (Claude Code, xhigh)35.91
6Claude Opus 5 (Claude Code, max)33.54
7GPT-5.6 Terra (Codex, max)33.46
8Claude Opus 5 (Claude Code, high)33.12
9GLM-5.3-Flash (OpenCode, max)32.42
10GPT-5.6 Luna (Codex, max)31.02

Interactive version: theaggregate.ai/benchmark?slug=pptbench · How It Works · Data refreshed daily, snapshot 2026-09-26.