DexHoldem Policy Bench - Scene-Preserving Success: leaderboard
Metric: Scene-preserving success rate (%; the primitive completes without disturbing other cards or chips, over 80 real-robot primitive trials (10 per pickup primitive, 5 per other primitive, 14 primitives) per policy, each policy trained as one multi-task policy on the 1,470 DexHoldem teleoperated demonstrations). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | π0.5 | 47.5 |
| 2 | π0 | 47.5 |
| 3 | RDT | 30 |
| 4 | DP (DINO) | 26.2 |
| 5 | DP-Transformer | 13.8 |
| 6 | RDT-small | 13.8 |
| 7 | ACT | 10 |
| 8 | BAKU | 6.2 |
| 9 | DP-UNet | 1.2 |
Interactive version: theaggregate.ai/benchmark?slug=dexholdem-policy-bench-scene-preserving-success · How It Works · Data refreshed daily, snapshot 2026-09-26.