ReLE - Language and Instruction Following: leaderboard
Metric: Accuracy (%). Source: nonelinear.com. 179 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek V4 Pro | 77 |
| 2 | GLM-5.3 Flash | 77 |
| 3 | Seed 2.0 Pro | 76 |
| 4 | GPT-5 | 75.6 |
| 5 | GPT-6 | 75.3 |
| 6 | GLM-5.3 | 75 |
| 7 | DeepSeek V3.2 (Thinking) | 74.7 |
| 8 | DeepSeek V4.1 Flash | 74.1 |
| 9 | Gemini 3.8 Flash | 73.5 |
| 10 | Kimi K2.7 Code | 73 |
| 11 | DeepSeek R1 0528 | 72.9 |
| 12 | Hunyuan-T1-20250711 | 72.9 |
| 13 | GPT-5.6 Pro Sol | 72.8 |
| 14 | Qwen 3.8 Max | 72.6 |
| 15 | Kimi K3 | 72.5 |
Interactive version: theaggregate.ai/benchmark?slug=rele-language-and-instruction-following · How It Works · Data refreshed daily, snapshot 2026-09-19.