K-BrowseComp: leaderboard
Metric: Pass@1 accuracy (%) on the 300-question K-BrowseComp-Verified subset; search_evals deep-research agent with Perplexity Search, 10 search calls per question, one run, GPT-5.4-mini answer extraction against the gold answer. Source: arxiv.org. Saturation forecast: Around January 2027. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GLM-5.1 | 30.67 |
| 2 | GPT-5.4 Mini | 30.67 |
| 3 | DeepSeek V4 Pro | 30 |
| 4 | Gemma 4 31B (IT) | 23.33 |
| 5 | Qwen 3.6 35B A3B | 12 |
| 6 | Gemini 3.1 Flash Lite | 11.33 |
| 7 | K-EXAONE | 10.33 |
| 8 | A.X 4.0 | 5.33 |
| 9 | HyperCLOVA X SEED Think (32B) | 2.33 |
Interactive version: theaggregate.ai/benchmark?slug=k-browsecomp · How It Works · Data refreshed daily, snapshot 2026-09-29.