Video-IFBench - Nested: leaderboard
Metric: Task-gated instruction satisfaction rate (%; TISR on nested instructions, where the model must follow the one true root-to-leaf path of a condition tree; checklist-based instruction following on 1.5K video samples with 39 semantic and format constraint types; a response counts only if it addresses the active task (the correct branch for conditional instructions) and satisfies every checklist item, judged by Qwen3.5-397B-A17B-Instruct or verified programmatically; Gemini and Doubao get video at 1 fps, GPT-5.4 50 frames, vision-only models frames plus timestamped subtitles). Source: arxiv.org. Saturation forecast: Around July 2027. 38 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Pro | 46 |
| 2 | Gemini 3 Flash | 30.7 |
| 3 | Qwen 3.5 397B A17B (Thinking) | 28.2 |
| 4 | Qwen 3.5 27B (Thinking) | 20 |
| 5 | Qwen 3.5 35B A3B (Thinking) | 19.8 |
| 6 | Qwen 3.5 122B A10B (Thinking) | 19.7 |
| 7 | Doubao-Seed-2.0-Pro-260215 | 14.4 |
| 8 | Qwen 3.5 9B (Thinking) | 11.9 |
| 9 | Qwen 3 VL 235B A22B (Thinking) | 10.3 |
| 10 | Gemma 4 31B (IT) | 10 |
| 11 | Qwen 3 VL 30B A3B Instruct | 9.5 |
| 12 | Qwen 3 VL 8B Instruct | 8.5 |
| 13 | Qwen 3.5 397B A17B (Non-reasoning) | 7.9 |
| 14 | GPT-5.4 | 7.6 |
| 15 | Qwen 3.5 27B (Non-reasoning) | 6.9 |
Interactive version: theaggregate.ai/benchmark?slug=video-ifbench-nested · How It Works · Data refreshed daily, snapshot 2026-09-26.