MemCalib: leaderboard

Metric: Sample Calibration Score (%; geometric-decay score over per-response memory over-use and under-use totals, DeepSeek-V4-Pro rubric judge, mean over 3 seeds; non-thinking mode). Source: arxiv.org. Saturation forecast: Around 2030. 9 models tracked.

Top models

#ModelScore
1GPT-5.6 Sol (Non-reasoning)46.25
2Claude Sonnet 4.636.44
3Kimi K2.6 (Non-reasoning)35.54
4GLM-5.2 (Non-reasoning)34.4
5DeepSeek V4 Flash (Non-reasoning)33.72
6Qwen 3 8B (Non-reasoning)31.17
7Qwen 3.5 35B A3B (Non-reasoning)26.54

Interactive version: theaggregate.ai/benchmark?slug=memcalib · How It Works · Data refreshed daily, snapshot 2026-09-26.