VEFX-Bench - Instruction Following: leaderboard
Metric: Instruction-following score as a coverage-adjusted model-level estimate (inverse-propensity-weighted mixed-effects) on the 1-4 rubric scale over the 300 curated VEFX-Bench source-video and instruction pairs, each system's edited video scored by the authors' VEFX-Reward-32B evaluator from soft expected predictions; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Kling o1 | 3.04 |
| 2 | Kling 3.0 Omni | 3.03 |
| 3 | Runway Gen-4.5 | 2.82 |
| 4 | Seedance 2.0 | 2.81 |
| 5 | Luma ray 3 | 2.7 |
| 6 | Grok Imagine | 2.61 |
| 7 | UniVideo | 2.29 |
| 8 | Luma ray 2 | 2.04 |
| 9 | VACE (VEFX-Bench checkpoint unspecified) | 2.03 |
| 10 | Wan 2.6 | 2.01 |
Interactive version: theaggregate.ai/benchmark?slug=vefx-bench-instruction-following · How It Works · Data refreshed daily, snapshot 2026-10-07.