WEval: leaderboard

Metric: Ranking correlation (Kendall-style pairwise sign agreement with the gold ranking, x100, -100 to 100, averaged over instructions); WEval writing reward-model test: 2,750 writing instructions with five requirements each; for every instruction DeepSeek-R1 answers the full instruction and versions with 1 to 4 requirements dropped, and the gold ranking orders the five responses by how many requirements were kept; LLMs are prompted to rank the responses, reward models score them; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 5 models tracked.

Top models

#ModelScore
1Qwen 2.5 72B Instruct93.6
2Qwen 2.5 7B Instruct85.4
3Skywork-Reward-V2-Qwen3-8B82.5
4Skywork-Reward-V2-Llama-3.1-8B75.4

Interactive version: theaggregate.ai/benchmark?slug=weval · How It Works · Data refreshed daily, snapshot 2026-10-07.