WEval - Prompt-Level: leaderboard
Metric: Prompt-level accuracy (%: share of instructions whose predicted ranking of the five responses equals the gold ranking exactly); WEval writing reward-model test: 2,750 writing instructions with five requirements each; for every instruction DeepSeek-R1 answers the full instruction and versions with 1 to 4 requirements dropped, and the gold ranking orders the five responses by how many requirements were kept; LLMs are prompted to rank the responses, reward models score them; higher is better. Source: arxiv.org. Saturation forecast: Around June 2028. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 2.5 72B Instruct | 78.7 |
| 2 | Qwen 2.5 7B Instruct | 60.9 |
| 3 | Skywork-Reward-V2-Qwen3-8B | 45.3 |
| 4 | Skywork-Reward-V2-Llama-3.1-8B | 32.9 |
Interactive version: theaggregate.ai/benchmark?slug=weval-prompt-level · How It Works · Data refreshed daily, snapshot 2026-10-07.