BrowseComp Long Context 128k: leaderboard
Long-context BrowseComp variant for hard-to-find web questions, evaluating whether models can use a 128k context budget to retrieve and reason over evidence.
Metric: Score (%). Source: llm-stats.com. Status: saturated. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.2 | 92 |
| 2 | GPT-5 | 90 |
| 3 | GPT-5.1 | 90 |
| 4 | GPT-5.1 (Thinking) | 90 |
| 5 | GPT-5.1 Instant | 90 |
Interactive version: theaggregate.ai/benchmark?slug=browsecomp-long-context-128k · How It Works · Data refreshed daily, snapshot 2026-09-05.