HELM Robo-Reward-Bench - Robo Reward Bench Progress Prediction: leaderboard

Metric: Absolute error. Source: crfm.stanford.edu. 18 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro (Preview 05-06)18.15
2GPT-4.1 (2025-04-14)20.46
3O1 (2024-12-17)22.2
4GPT-4o (2024-11-20)22.89
5GPT-4.524.67
6Gemini 2.5 Flash (Preview 04-17)25.1
7Gemini 2.5 Flash (Preview 05-20)26.56
8Gemini 2.0 Flash Lite27.54
9Gemini 2.0 Flash27.63
10GPT-4.1 Mini36.17

Interactive version: theaggregate.ai/benchmark?slug=helm-robo-reward-bench-robo-reward-bench-progress-prediction · How It Works · Data refreshed daily, snapshot 2026-09-08.