OSCBench - Compositional Scenarios: leaderboard
Metric: Object state change score (mean of accuracy and consistency) on OSCBench's 12 compositional scenarios (chained actions on one object), human ratings (three raters, 1-5 Likert, printed on a 0-1 scale, x100), one generated video per scenario; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 6 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Veo-3.1-Fast | 80.5 | |
| 2 | Kling 2.5 Turbo | 69.9 | |
| 3 | Wan2.2 | 59.4 | |
| 4 | HunyuanVideo-1.5 | 55.6 | |
| 5 | HunyuanVideo | 43.7 | |
| 6 | Open-Sora-2.0 | 41.6 |
Interactive version: theaggregate.ai/benchmark?slug=oscbench-compositional-scenarios · How It Works · Data refreshed daily, snapshot 2026-10-11.