WEval - Prompt-Level: leaderboard

Metric: Prompt-level accuracy (%: share of instructions whose predicted ranking of the five responses equals the gold ranking exactly); WEval writing reward-model test: 2,750 writing instructions with five requirements each; for every instruction DeepSeek-R1 answers the full instruction and versions with 1 to 4 requirements dropped, and the gold ranking orders the five responses by how many requirements were kept; LLMs are prompted to rank the responses, reward models score them; higher is better. Source: arxiv.org. Saturation forecast: Around June 2028. 5 models tracked.

Top models

#ModelScore
1Qwen 2.5 72B Instruct78.7
2Qwen 2.5 7B Instruct60.9
3Skywork-Reward-V2-Qwen3-8B45.3
4Skywork-Reward-V2-Llama-3.1-8B32.9

Interactive version: theaggregate.ai/benchmark?slug=weval-prompt-level · How It Works · Data refreshed daily, snapshot 2026-10-07.