HELM Robo-Reward-Bench - Robo Reward Bench Discrete: leaderboard
Metric: Absolute error. Source: crfm.stanford.edu. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4.1 (2025-04-14) | 1.4 |
| 2 | GPT-4o (2024-11-20) | 1.44 |
| 3 | Gemini 2.0 Flash | 1.66 |
Interactive version: theaggregate.ai/benchmark?slug=helm-robo-reward-bench-robo-reward-bench-discrete · How It Works · Data refreshed daily, snapshot 2026-09-08.