BrowseComp Long Context 128k: leaderboard

Long-context BrowseComp variant for hard-to-find web questions, evaluating whether models can use a 128k context budget to retrieve and reason over evidence.

Metric: Score (%). Source: llm-stats.com. Status: saturated. 5 models tracked.

Top models

#ModelScore
1GPT-5.292
2GPT-590
3GPT-5.190
4GPT-5.1 (Thinking)90
5GPT-5.1 Instant90

Interactive version: theaggregate.ai/benchmark?slug=browsecomp-long-context-128k · How It Works · Data refreshed daily, snapshot 2026-09-05.