HMMT February 2026: leaderboard

Official Hugging Face benchmark for model performance on the February 2026 Harvard-MIT Mathematics Tournament problem set.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 25 models tracked.

Top models

#ModelScore
1Qwen 3.7 Max97.1
2GPT-5.596.7
3Claude Opus 4.896.7
4Claude Opus 4.6 (Max)96.2
5DeepSeek V4 Pro95.2
6Qwen 3.7 Plus92.9
7Kimi K2.6 (Thinking)92.7
8GLM-5.292.5
9Inkling Small90.2
10Qwen 3.5 397B A17B87.88
11Qwen 3.6 Plus87.8
12Gemini 3.1 Pro (Preview)87.3
13Kimi K2.587.12
14GLM-586.36
15Step 3.5 Flash86.36

Interactive version: theaggregate.ai/benchmark?slug=hmmt-february-2026 · How It Works · Data refreshed daily, snapshot 2026-09-05.