MemCalib - Exact Calibration: leaderboard

Metric: Exact calibration (%; share of responses with zero memory over-use and zero under-use across all atoms, DeepSeek-V4-Pro rubric judge, mean over 3 seeds; non-thinking mode). Source: arxiv.org. Saturation forecast: Around 2030. 9 models tracked.

Top models

#ModelScore
1GPT-5.6 Sol (Non-reasoning)28.4
2Kimi K2.6 (Non-reasoning)20
3Claude Sonnet 4.619.09
4GLM-5.2 (Non-reasoning)18.2
5DeepSeek V4 Flash (Non-reasoning)16.96
6Qwen 3 8B (Non-reasoning)15.29
7Qwen 3.5 35B A3B (Non-reasoning)12.24

Interactive version: theaggregate.ai/benchmark?slug=memcalib-exact-calibration · How It Works · Data refreshed daily, snapshot 2026-09-26.