K-BrowseComp - Synthetic: leaderboard

Metric: Pass@1 accuracy (%) on the 100-question synthetic split (agent-generated Korean browsing questions kept only when gpt-5.4-mini and gemini-3-flash-preview failed them); same harness and grading. Source: arxiv.org. Saturation forecast: Around February 2028. 11 models tracked.

Top models

#ModelScore
1GPT-5.526
2DeepSeek V4 Pro22
3GLM-5.119
4Gemma 4 31B (IT)17
5Qwen 3.6 35B A3B15
6K-EXAONE13
7Gemini 3.1 Flash Lite11
8HyperCLOVA X SEED Think (32B)2
9A.X 4.01

Interactive version: theaggregate.ai/benchmark?slug=k-browsecomp-synthetic · How It Works · Data refreshed daily, snapshot 2026-09-29.