ComBench: leaderboard

Metric: Overall average normalized score (%) over four sampled solutions across the 100 problems (proof rubric score for analysis items, verifier-gated score for construction items); higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.

Top models

#ModelScore
1GPT-5.565.4
2Gemini 3.1 Pro (Preview) (Thinking)60.3
3Kimi K2.653.5
4DeepSeek V4 Pro45.2
5Qwen 3.6 Max Preview24.9
6GLM-5.123.6
7Qwen 3.6 35B A3B20.3
8Nemotron Cascade 2 30B A3B19.6
9Gemma 4 31B (IT)16.8

Interactive version: theaggregate.ai/benchmark?slug=combench · How It Works · Data refreshed daily, snapshot 2026-09-29.