SpatiaLQA - Precondition F1: leaderboard

Metric: Precondition F1 (%, 0-100): harmonic mean of the pooled recall and precision of the predicted step dependencies between matched steps; SpatiaLQA: 9,605 questions on 2,401 photographs of 241 real indoor scenes (13 categories), each asking for the ordered steps of a multi-step task and every step's preconditions in a fixed JSON template; GPT-4o matches predicted to annotated steps per image (a 0/1 matrix filtered by the Hungarian algorithm), then recall and precision are pooled over all samples; open models with their Hugging Face recommended settings, proprietary models at temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 40 models tracked.

Top models

#ModelScoreOverall rank
1GPT-547#91
2GPT-5 Mini41.2#176
3Gemini 2.5 Pro39.3#145
4Gemini 2.5 Flash38.9#237
5GPT-4.138#240
6Claude Sonnet 436.3#194
7Qwen 2.5 VL 72B Instruct31.2#364
8GPT-4.1 Mini30#346
9Claude 3.7 Sonnet27.9#241
10GPT-4o25.1#333
11Qwen 3 VL 4B Instruct19.6#506
12Qwen 3 VL 8B Instruct19#401
13Pixtral-12B19#795
14GLM-4.1V-9B (Thinking)18.6#457 (GLM-4.1V-9B)
15GPT-4o Mini18.4#588

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=spatialqa-precondition-f1 · How It Works · Data refreshed daily, snapshot 2026-10-11.