Flat-Pack Bench - Temporal Localization: leaderboard

Metric: TLoc: identify the events just before or after the state shown in the visual prompt, 103 questions; multiple-choice accuracy (%, regex exact match) on furniture-assembly videos from IMaW with one or two segmented visual-prompt frames, zero-shot with greedy decoding (GPT-5 at its default); each row is the model's best of up to six settings (key-frame or trimmed video, mixed-media, collage or concat visual prompt) chosen by overall micro-average accuracy, the paper's main-table protocol; higher is better. Source: arxiv.org. Saturation forecast: Around July 2027. 30 models tracked.

Top models

#ModelScore
1GPT-553.4
2Qwen 3 VL 32B Instruct46.6
3Gemini 2.5 Pro44.66
4Gemini 3.1 Pro (Preview)43.69
5Gemini 2.5 Flash41.75
6InternVL3-78B39.81
7InternVL3-38B37.86
8Qwen 3 VL 4B Instruct33.01
9Qwen 3 VL 8B (Thinking)33.01
10Qwen 3 VL 8B Instruct30.1
11Qwen 2.5 VL 72B Instruct30.1
12Qwen 2.5 VL 32B Instruct29.13
13Qwen 3 VL 235B A22B Instruct25.24
14Qwen 3 VL 4B (Thinking)25.24
15Qwen 3 VL 30B A3B Instruct22.33

Interactive version: theaggregate.ai/benchmark?slug=flat-pack-bench-temporal-localization · How It Works · Data refreshed daily, snapshot 2026-10-07.