PlanarBench: leaderboard

Metric: Score (out of 199): sum over the 199 planar graphs (2-7 vertices) of the per-graph drawing score, 1 for a strictly valid straight-line crossing-free ASCII drawing, otherwise half credit each for a coordinate-valid node placement and for BFS-valid edge connectivity; one attempt per graph, single fixed prompt; the paper prints its top 21 of 91 configurations; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 21 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)151
2Claude Opus 4.5 (Thinking)144
3GPT-5 Codex130
4GPT-5 Codex (High)127
5GPT-5124.5
6GPT-5 (High)124
7Claude Sonnet 4.5 (Thinking)112.5
8GPT-5 Mini (High)111
9GPT-5.2 (Low)101.5
10Gemini 2.5 Pro (Preview 05-06)96
11Claude 3.7 Sonnet (Thinking)95.5
12O4 Mini93.5
13O191
14Gemini 2.5 Pro (Preview 03-25)85
15GLM-4.784

Interactive version: theaggregate.ai/benchmark?slug=planarbench · How It Works · Data refreshed daily, snapshot 2026-09-29.