GRIPS-hard — leaderboard

Metric: Accuracy (%). Source: huggingface.co. 27 models tracked.

Top models

#ModelScore
1GPT-5.5 (xHigh)87.3
2GPT-5.5 (High)85.3
3GPT-5.580.7
4Claude Opus 4.8 (High)76
5Gemini 3.5 Flash (High)74
6GPT-5.5 (Low)70.7
7DeepSeek V4 Pro (Reasoning)67.3
8Claude Sonnet 4.6 (High)60.4
9GLM-5.2 (Reasoning)54
10DeepSeek V4 Flash (Reasoning)49.3
11GPT-5.4 Mini (High)42
12Mistral Medium 3.5 (High)26.2
13GPT-5.4 Mini17.3
14Claude Haiku 4.514
15Mistral Small14

Interactive version: theaggregate.ai/benchmark?slug=grips-hard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.