ComBench - Analysis-Centric: leaderboard

Metric: Average normalized proof score (%) on the 50 analysis-centric problems; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 10 models tracked.

Top models

#ModelScore
1GPT-5.562.4
2Gemini 3.1 Pro (Preview) (Thinking)56.1
3Kimi K2.643.5
4DeepSeek V4 Pro37.8
5Nemotron Cascade 2 30B A3B21.8
6GLM-5.121.6
7Qwen 3.6 Max Preview21.4
8Qwen 3.6 35B A3B17.9
9Gemma 4 31B (IT)16.1

Interactive version: theaggregate.ai/benchmark?slug=combench-analysis-centric · How It Works · Data refreshed daily, snapshot 2026-09-29.