ReFigBench (Direct Code Generation): leaderboard

Metric: Overall rubric score (0-100) under the direct workflow: the agent writes a program that builds the slide with a general PPTX library (python-pptx); GPT-5.4 judge (gpt-5.4-2026-03-05, low effort, first of three runs) scores the rendered slide and its native object tree against the source figure on text 20, semantic structure 30, layout 15, editability 25 and visual detail 10 points; unrenderable artifacts score 0 and a full-slide raster paste is capped at 50; mean over 1,000 arXiv overview figures. Source: arxiv.org. Saturation forecast: Around December 2026. 5 models tracked.

Top models

#ModelScore
1GPT-5.5 (Codex, xHigh)74.2
2Claude Opus 4.6 (Claude Code)69.4

Interactive version: theaggregate.ai/benchmark?slug=refigbench-direct-code-generation · How It Works · Data refreshed daily, snapshot 2026-09-26.