WuYuEval - Expert Module - Elo: leaderboard
Metric: Elo rating from pairwise judged comparisons. Source: arxiv.org. 33 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.5 | 2022.8 |
| 2 | Kimi K2 (Thinking) | 1954.97 |
| 3 | GPT-5.2 | 1904.58 |
| 4 | DeepSeek V3.2 (Non-reasoning) | 1788.91 |
| 5 | GLM-4.6 (Non-reasoning) | 1755.15 |
| 6 | Qwen 3 235B A22B Instruct | 1723.93 |
| 7 | Qwen 3 235B A22B (Thinking) | 1723.74 |
| 8 | Gemini 3 Pro (Preview) | 1719.25 |
| 9 | DeepSeek V3.2 (Thinking) | 1703.23 |
| 10 | Qwen 3 Max | 1608.84 |
| 11 | GPT-OSS-20B | 1512.59 |
| 12 | Gemma 3 12B (IT) | 1498.67 |
| 13 | Gemma 3 27B (IT) | 1458.88 |
| 14 | GLM-4.6 (Thinking) | 1423.8 |
| 15 | Qwen 3 32B (Thinking) | 1399.3 |
Interactive version: theaggregate.ai/benchmark?slug=wuyueval-expert-module-elo · How It Works · Data refreshed daily, snapshot 2026-09-19.