PlanarBench: leaderboard
Metric: Score (out of 199): sum over the 199 planar graphs (2-7 vertices) of the per-graph drawing score, 1 for a strictly valid straight-line crossing-free ASCII drawing, otherwise half credit each for a coordinate-valid node placement and for BFS-valid edge connectivity; one attempt per graph, single fixed prompt; the paper prints its top 21 of 91 configurations; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 21 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Pro (Preview) | 151 |
| 2 | Claude Opus 4.5 (Thinking) | 144 |
| 3 | GPT-5 Codex | 130 |
| 4 | GPT-5 Codex (High) | 127 |
| 5 | GPT-5 | 124.5 |
| 6 | GPT-5 (High) | 124 |
| 7 | Claude Sonnet 4.5 (Thinking) | 112.5 |
| 8 | GPT-5 Mini (High) | 111 |
| 9 | GPT-5.2 (Low) | 101.5 |
| 10 | Gemini 2.5 Pro (Preview 05-06) | 96 |
| 11 | Claude 3.7 Sonnet (Thinking) | 95.5 |
| 12 | O4 Mini | 93.5 |
| 13 | O1 | 91 |
| 14 | Gemini 2.5 Pro (Preview 03-25) | 85 |
| 15 | GLM-4.7 | 84 |
Interactive version: theaggregate.ai/benchmark?slug=planarbench · How It Works · Data refreshed daily, snapshot 2026-09-29.