TraversalBench - Reading Order: leaderboard
Metric: Token-level accuracy (%): position-wise share of correctly ordered vertex labels after standardized parsing, when a model recovers the vertex sequence of a polyline drawn in one of four scan-order regimes (left-to-right or right-to-left then top-to-bottom, top-to-bottom then left-to-right or right-to-left), macro-averaged over lane-count and vertex-count combinations within each regime; shared prompt template and output format; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 | 98.7 |
| 2 | Gemini 3.1 Flash Lite | 90.1 |
| 3 | Claude Opus 4.6 | 63.9 |
Interactive version: theaggregate.ai/benchmark?slug=traversalbench-reading-order · How It Works · Data refreshed daily, snapshot 2026-10-07.