Graphwalks BFS >128k: leaderboard
Long-context Graphwalks breadth-first-search task over 128k tokens, stressing graph traversal when the relevant edges are spread through a very large input.
Metric: Score (%). Source: llm-stats.com. Status: saturation imminent. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.6 Sol | 90.7 |
| 2 | GPT-5.6 Luna | 81.3 |
| 3 | Claude Mythos Preview | 80 |
| 4 | GPT-5.6 Terra | 76.9 |
| 5 | Claude Opus 4.8 | 68.1 |
| 6 | Claude Opus 4.6 | 61.5 |
| 7 | GPT-5.5 | 45.4 |
| 8 | GPT-5.4 | 21.4 |
| 9 | GPT-4.1 | 19 |
| 10 | GPT-4.1 Mini | 15 |
| 11 | GPT-4.1 Nano | 2.9 |
Interactive version: theaggregate.ai/benchmark?slug=graphwalks-bfs-over-128k · How It Works · Data refreshed daily, snapshot 2026-09-05.