WEval: leaderboard
Metric: Ranking correlation (Kendall-style pairwise sign agreement with the gold ranking, x100, -100 to 100, averaged over instructions); WEval writing reward-model test: 2,750 writing instructions with five requirements each; for every instruction DeepSeek-R1 answers the full instruction and versions with 1 to 4 requirements dropped, and the gold ranking orders the five responses by how many requirements were kept; LLMs are prompted to rank the responses, reward models score them; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 2.5 72B Instruct | 93.6 |
| 2 | Qwen 2.5 7B Instruct | 85.4 |
| 3 | Skywork-Reward-V2-Qwen3-8B | 82.5 |
| 4 | Skywork-Reward-V2-Llama-3.1-8B | 75.4 |
Interactive version: theaggregate.ai/benchmark?slug=weval · How It Works · Data refreshed daily, snapshot 2026-10-07.