Graphwalks BFS 256k F1 — leaderboard

Graphwalks breadth-first-search long-context reasoning task reported at 256k context with F1 scoring.

Metric: F1 (self-reported). Source: benchmarklist.com. Status: saturation imminent. 7 models tracked.

Top models

#ModelScore
1Claude Mythos 591.1
2Claude Opus 4.885.9
3Claude Mythos Preview85.7
4Claude Opus 4.776.9
5GPT-5.573.7
6GPT-5.462.5
7Claude Opus 4.661.1

Interactive version: theaggregate.ai/benchmark?slug=graphwalks-bfs-256k-f1 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.