Video-IFBench - Nested: leaderboard

Metric: Task-gated instruction satisfaction rate (%; TISR on nested instructions, where the model must follow the one true root-to-leaf path of a condition tree; checklist-based instruction following on 1.5K video samples with 39 semantic and format constraint types; a response counts only if it addresses the active task (the correct branch for conditional instructions) and satisfies every checklist item, judged by Qwen3.5-397B-A17B-Instruct or verified programmatically; Gemini and Doubao get video at 1 fps, GPT-5.4 50 frames, vision-only models frames plus timestamped subtitles). Source: arxiv.org. Saturation forecast: Around July 2027. 38 models tracked.

Top models

#ModelScore
1Gemini 3 Pro46
2Gemini 3 Flash30.7
3Qwen 3.5 397B A17B (Thinking)28.2
4Qwen 3.5 27B (Thinking)20
5Qwen 3.5 35B A3B (Thinking)19.8
6Qwen 3.5 122B A10B (Thinking)19.7
7Doubao-Seed-2.0-Pro-26021514.4
8Qwen 3.5 9B (Thinking)11.9
9Qwen 3 VL 235B A22B (Thinking)10.3
10Gemma 4 31B (IT)10
11Qwen 3 VL 30B A3B Instruct9.5
12Qwen 3 VL 8B Instruct8.5
13Qwen 3.5 397B A17B (Non-reasoning)7.9
14GPT-5.47.6
15Qwen 3.5 27B (Non-reasoning)6.9

Interactive version: theaggregate.ai/benchmark?slug=video-ifbench-nested · How It Works · Data refreshed daily, snapshot 2026-09-26.