NeoEvalPlusN — leaderboard
Public leaderboard for proprietary command-following, distractor-resistance, expectation-breaking, poem, and stylized-writing tests run mainly on open-source LLM variants.
Metric: Mean Score (0-6). Source: huggingface.co. Status: years away from saturation. 203 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek V3 | 3.7 |
| 2 | Kimi K2 0905 | 3.2 |
| 3 | Kimi K2 | 3.1 |
| 4 | Llama 4 Scout Instruct | 2.05 |
| 5 | Gemma 3 27B (IT) | 2 |
| 6 | Gemma 2 27B (IT) | 2 |
| 7 | Gemma 2 9B (IT) | 1.95 |
Interactive version: theaggregate.ai/benchmark?slug=neoevalplusn · How the rankings work · Data refreshed daily, snapshot 2026-07-22.