NeoEvalPlusN — leaderboard

Public leaderboard for proprietary command-following, distractor-resistance, expectation-breaking, poem, and stylized-writing tests run mainly on open-source LLM variants.

Metric: Mean Score (0-6). Source: huggingface.co. Status: years away from saturation. 203 models tracked.

Top models

#ModelScore
1DeepSeek V33.7
2Kimi K2 09053.2
3Kimi K23.1
4Llama 4 Scout Instruct2.05
5Gemma 3 27B (IT)2
6Gemma 2 27B (IT)2
7Gemma 2 9B (IT)1.95

Interactive version: theaggregate.ai/benchmark?slug=neoevalplusn · How the rankings work · Data refreshed daily, snapshot 2026-07-22.