RoboTrustBench - Constraint-Sensitive: leaderboard

Metric: Overall average (0-100) of the five dimension means over 12 human-rated criteria on the Constraint-Sensitive (feasible but ambiguous, occluded, cluttered or trajectory-constrained instruction) scenario, stratified human-evaluation subset (10 instruction-image pairs from each of the 18 fine-grained subcategories, 180 in all; three raters score each generated video 1 to 5 per criterion, normalized to 0-1 and printed x100 here): an image-to-video world model generates a robot-arm manipulation video from a real DROID initial frame and instruction, rated for scene entity alignment, spatiotemporal consistency, interaction rationality, task execution and visual quality; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 7 models tracked.

Top models

#ModelScore
1Kling 2.675.9
2Veo 3.1 Fast65.7
3Cosmos-Predict2.5-14B61.2
4Wan2.2-I2V-A14B58.6
5LingBot-World57.8
6Cosmos-Predict2.5-2B57
7HunyuanVideo-1.550.8

Interactive version: theaggregate.ai/benchmark?slug=robotrustbench-constraint-sensitive · How It Works · Data refreshed daily, snapshot 2026-09-29.