Flat-Pack Bench - Temporal Ordering: leaderboard

Metric: TOrd: order the connection events of the highlighted parts, 155 questions; multiple-choice accuracy (%, regex exact match) on furniture-assembly videos from IMaW with one or two segmented visual-prompt frames, zero-shot with greedy decoding (GPT-5 at its default); each row is the model's best of up to six settings (key-frame or trimmed video, mixed-media, collage or concat visual prompt) chosen by overall micro-average accuracy, the paper's main-table protocol; higher is better. Source: arxiv.org. Saturation forecast: Around 2033. 30 models tracked.

Top models

#ModelScore
1InternVL3-78B43.87
2InternVL3-38B42.58
3InternVL3-14B42.58
4Qwen 2.5 VL 72B Instruct41.29
5Gemini 2.5 Pro40.65
6GPT-540.65
7Qwen 3 VL 32B Instruct38.71
8Qwen 3 VL 32B (Thinking)38.71
9Qwen 3 VL 235B A22B Instruct37.42
10Qwen 3 VL 8B Instruct36.13
11Gemini 3.1 Pro (Preview)34.84
12Qwen 2.5 VL 32B Instruct34.84
13Qwen 3 VL 4B Instruct34.19
14Qwen 3 VL 8B (Thinking)34.19
15Gemini 2.5 Flash31.61

Interactive version: theaggregate.ai/benchmark?slug=flat-pack-bench-temporal-ordering · How It Works · Data refreshed daily, snapshot 2026-10-07.