K-BrowseComp - Calibration Error: leaderboard

Metric: Expected calibration error (%) of the stated confidence on K-BrowseComp-Verified (five equal-width bins, weighted gap between mean confidence and accuracy); lower is better. Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.

Top models

#ModelScore
1DeepSeek V4 Pro17.72
2Gemma 4 31B (IT)23.66
3K-EXAONE24.09
4GLM-5.127.07
5GPT-5.531.86
6GPT-5.4 Mini37.88
7Qwen 3.6 35B A3B47.89
8A.X 4.047.89
9Gemini 3.1 Flash Lite56.55
10HyperCLOVA X SEED Think (32B)77.37

Interactive version: theaggregate.ai/benchmark?slug=k-browsecomp-calibration-error · How It Works · Data refreshed daily, snapshot 2026-09-29.