Long-form RewardBench - Reasoning: leaderboard

Metric: Pairwise accuracy (%) on the math reasoning subset of Long-form RewardBench: verified correctness decides the chosen response and the reward model must prefer it over an incorrect one; sequence classifiers score each response, generative judges select the better one; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 22 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Flash (Preview 04-17)90.2#230
2Claude Opus 4 (20250514)80.1#148
3GPT-4.178.4#240
4Claude 3.7 Sonnet (20250219)73.4#196
5GPT-4o (2024-08-06)73.1#326
6Claude 3.5 Sonnet (20240620)70.3#293

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=long-form-rewardbench-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-11.