RoboTrustBench - Normal (GPT-5.4 Judge): leaderboard

Metric: Overall average (0-100) of the five dimension means over the same 12 criteria scored by a GPT-5.4 judge on the Normal (feasible, unambiguous instruction) scenario, automatic evaluation that the paper runs to scale to the full benchmark (533 pairs in this scenario; the judge sees the instruction, the initial image and 20 uniformly sampled frames and cites frame evidence before scoring each criterion 1 to 5, normalized to 0-1 and printed x100 here): an image-to-video world model generates a robot-arm manipulation video from a real DROID initial frame and instruction, rated for scene entity alignment, spatiotemporal consistency, interaction rationality, task execution and visual quality; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 7 models tracked.

Top models

#ModelScore
1Kling 2.688.6
2Veo 3.1 Fast85.6
3Cosmos-Predict2.5-14B83.8
4Cosmos-Predict2.5-2B83.7
5LingBot-World82.1
6Wan2.2-I2V-A14B81.7
7HunyuanVideo-1.580.6

Interactive version: theaggregate.ai/benchmark?slug=robotrustbench-normal-gpt-5-4-judge · How It Works · Data refreshed daily, snapshot 2026-09-29.