GIM — leaderboard

Grounded Integration Measure from Meta FAIR: 820 multimodal and text-grounded problems requiring integrated reasoning across quantitative, spatial, language, world-knowledge, and document tasks.

Metric: IRT ability (theta). Source: arxiv.org. Status: saturation imminent. 46 models tracked.

Top models

#ModelScore
1GPT-5.4 Pro (xHigh)2.16
2Gemini 3.1 Pro (Preview) (High)2.04
3GPT-5.4 (xHigh)1.64
4GPT-5.4 (High)1.59
5Claude Opus 4.7 (High)1.53
6Claude Opus 4.6 (High)1.48
7Gemini 3.1 Pro (Preview) (Low)1.45
8GPT-5.4 (Medium)1.39
9Claude Sonnet 4.6 (High)1.12
10Claude Opus 4.6 (Medium)1.09
11GPT-5 (High)0.9
12GPT-5.4 (Low)0.82
13GPT-5 (Medium)0.59
14GPT-5 (Low)0.05
15GPT-5.4 Mini (High)0.02

Interactive version: theaggregate.ai/benchmark?slug=gim · How the rankings work · Data refreshed daily, snapshot 2026-07-22.