General365 - Branching & Enumeration: leaderboard

Metric: Accuracy (%) on the problems labelled Branching & Enumeration (a problem may carry several challenge labels), out of 1,460 problems (365 hand-written seeds and 1,095 expansions); numeric answers checked with math-verify, choice and text answers graded by GPT-4.1; temperature 1.0 for reasoning and 0.7 for chat models, highest available reasoning effort; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 26 models tracked.

Top models

#ModelScore
1GLM-5 (Thinking)69.1
2GPT-5 (High)65.4
3Qwen 3 Max (Thinking)65.3
4Gemini 3 Flash64.5
5Gemini 3 Pro64.3
6Kimi K2.5 (Thinking)64.3
7GLM-4.7 (Thinking)63.8
8DeepSeek V3.2 Speciale63.6
9GPT-5.1 (High)62.9
10DeepSeek V3.2 (Thinking)62.5
11Qwen 3.5 397B A17B (Thinking)62.1
12Kimi K2 (Thinking)61.9
13LongCat-Flash-Thinking-260161.4
14Grok 4.1 Fast (Reasoning)59.9
15GLM-4.6 (Thinking)59.9

Interactive version: theaggregate.ai/benchmark?slug=general365-branching-enumeration · How It Works · Data refreshed daily, snapshot 2026-10-07.