ReFigBench (Direct Code Generation) - Human Preference Elo: leaderboard
Metric: Human preference Elo (direct workflow: the agent writes a program that builds the slide with a general PPTX library (python-pptx); Bradley-Terry strength on the Elo scale (mean 1,000 over the ten configurations) from 4,270 blinded pairwise human comparisons of renderings of the same source figure, a tie counting half a win). Source: arxiv.org. Saturation forecast: Rough model projection: around 2027. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.5 (Codex, xHigh) | 1168.8 |
| 2 | Claude Opus 4.6 (Claude Code) | 885.5 |
Interactive version: theaggregate.ai/benchmark?slug=refigbench-direct-code-generation-human-preference-elo · How It Works · Data refreshed daily, snapshot 2026-09-26.