ReFigBench (Direct Code Generation) - Human Preference Elo: leaderboard

Metric: Human preference Elo (direct workflow: the agent writes a program that builds the slide with a general PPTX library (python-pptx); Bradley-Terry strength on the Elo scale (mean 1,000 over the ten configurations) from 4,270 blinded pairwise human comparisons of renderings of the same source figure, a tie counting half a win). Source: arxiv.org. Saturation forecast: Rough model projection: around 2027. 5 models tracked.

Top models

#ModelScore
1GPT-5.5 (Codex, xHigh)1168.8
2Claude Opus 4.6 (Claude Code)885.5

Interactive version: theaggregate.ai/benchmark?slug=refigbench-direct-code-generation-human-preference-elo · How It Works · Data refreshed daily, snapshot 2026-09-26.