RMGAP - Best-of-4: leaderboard

Metric: Best-of-4 accuracy (%; chosen response scored above all three alternatives, mean of chat, writing, reasoning and safety; scalar reward models score each response, DPO models by their implicit reward against the reference model). Source: arxiv.org. Saturation forecast: Estimated already saturated. 10 models tracked.

Top models

#ModelScore
1Skywork-Reward-V2-Qwen3-8B49.16
2URM-LLaMa-3.1-8B45.28
3InternLM2-7B-Reward38.23
4Llama-3.1-Tulu-3-8B-DPO31.07
5zephyr-7B-beta29.91

Interactive version: theaggregate.ai/benchmark?slug=rmgap-best-of-4 · How It Works · Data refreshed daily, snapshot 2026-09-26.