OSCBench - Regular Scenarios: leaderboard
Metric: Object state change score (mean of accuracy and consistency) on OSCBench's 108 regular scenarios (common cooking action-object pairs), human ratings (three raters, 1-5 Likert, printed on a 0-1 scale, x100), one generated video per scenario; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 6 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Veo-3.1-Fast | 79.7 | |
| 2 | Kling 2.5 Turbo | 74.4 | |
| 3 | Wan2.2 | 63.5 | |
| 4 | HunyuanVideo-1.5 | 57.2 | |
| 5 | HunyuanVideo | 47.2 | |
| 6 | Open-Sora-2.0 | 41 |
Interactive version: theaggregate.ai/benchmark?slug=oscbench-regular-scenarios · How It Works · Data refreshed daily, snapshot 2026-10-11.