H2R-Bench (Video-Conditioned) - Parallel-Jaw Gripper: leaderboard
Metric: H2RCore (0-100) = 100 (0.15 goal-state completion + 0.15 action-event completion + 0.30 functional contact transfer + 0.30 embodiment correctness + 0.10 video quality); the first four are 0-4 evidence ratings by Gemini 3.5 Flash, Qwen3.7-Plus and GPT-5.4 on 25 sampled frames, averaged, and video quality combines MUSIQ, CLIP aesthetic, temporal-stability and AMT-S interpolation scores; 120 egocentric human manipulation videos in six families, each transformed into a robot manipulation video; the model receives the full source video; target embodiment: a parallel-jaw gripper. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Seedance 2.0 | 77.3 |
| 2 | Wan2.7 | 76.5 |
| 3 | Kling-V3 | 74.5 |
Interactive version: theaggregate.ai/benchmark?slug=h2r-bench-video-conditioned-parallel-jaw-gripper · How It Works · Data refreshed daily, snapshot 2026-09-29.