Graphwalks BFS 256k F1: leaderboard

Graphwalks breadth-first-search long-context reasoning task reported at 256k context with F1 scoring.

Metric: F1 (self-reported). Source: benchmarklist.com. Status: years away from saturation. 10 models tracked.

Top models

#ModelScore
1GPT-5.6 Sol95.4
2GPT-5.6 Terra94.9
3GPT-5.6 Luna92.5
4Claude Mythos 591.1
5GPT-5.588.1
6Claude Opus 4.885.9
7Claude Mythos Preview85.7
8Claude Opus 4.776.9
9GPT-5.462.5
10Claude Opus 4.661.1

Interactive version: theaggregate.ai/benchmark?slug=graphwalks-bfs-256k-f1 · How It Works · Data refreshed daily, snapshot 2026-09-05.