HMMT February 2026 — leaderboard

Official Hugging Face benchmark for model performance on the February 2026 Harvard-MIT Mathematics Tournament problem set.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturated. 24 models tracked.

Top models

#ModelScore
1Qwen 3.7 Max (Max)97.1
2GPT-5.596.7
3Claude Opus 4.896.7
4Claude Opus 4.6 (Max)96.2
5DeepSeek V4 Pro95.2
6Qwen 3.7 Plus92.9
7GLM-5.292.5
8Qwen 3.5 397B A17B87.88
9Qwen 3.6 Plus87.8
10Gemini 3.1 Pro (Preview)87.3
11GLM-586.36
12Step 3.5 Flash86.36
13Nemotron 3 Super84.85
14MiniMax-M384.4
15DeepSeek V3.284.09

Interactive version: theaggregate.ai/benchmark?slug=hmmt-february-2026 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.