SpatiaLQA - Content F1: leaderboard

Metric: Step-content F1 (%, 0-100): harmonic mean of the pooled recall and precision of matched step contents; SpatiaLQA: 9,605 questions on 2,401 photographs of 241 real indoor scenes (13 categories), each asking for the ordered steps of a multi-step task and every step's preconditions in a fixed JSON template; GPT-4o matches predicted to annotated steps per image (a 0/1 matrix filtered by the Hungarian algorithm), then recall and precision are pooled over all samples; open models with their Hugging Face recommended settings, proprietary models at temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 41 models tracked.

Top models

#ModelScoreOverall rank
1GPT-576#91
2Gemini 2.5 Pro74.3#145
3GPT-5 Mini73.9#176
4GPT-4.173.5#240
5Gemini 2.5 Flash72.9#237
6Qwen 2.5 VL 72B Instruct69.6#364
7Claude Sonnet 468.9#194
8GPT-4o67.4#333
9GPT-4.1 Mini66.5#346
10GPT-4o Mini63.5#588
11Claude 3.7 Sonnet59.3#241
12GLM-4.1V-9B (Thinking)57.5#457 (GLM-4.1V-9B)
13Pixtral-12B57#795
14Qwen 2.5 VL 32B Instruct54#443
15Qwen 3 VL 8B Instruct52.7#401

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=spatialqa-content-f1 · How It Works · Data refreshed daily, snapshot 2026-10-11.