Graphwalks BFS 1M F1: leaderboard

Graphwalks breadth-first-search long-context reasoning task reported at 1M context with F1 scoring.

Metric: F1 (self-reported). Source: benchmarklist.com. Status: years away from saturation. 10 models tracked.

Top models

#ModelScore
1GPT-5.6 Sol83.4
2Claude Mythos 579.4
3GPT-5.6 Terra75.5
4Claude Mythos Preview74.3
5Claude Opus 4.868.1
6GPT-5.557.1
7GPT-5.6 Luna56.8
8Claude Opus 4.740.3
9Claude Opus 4.616.3
10GPT-5.49.4

Interactive version: theaggregate.ai/benchmark?slug=graphwalks-bfs-1m-f1 · How It Works · Data refreshed daily, snapshot 2026-09-05.