NCP-Bench: leaderboard
Metric: Average turns until termination (out of 100): mean over the environments of the turns played before the first confirmed conflict, the 100-turn limit or full commitment satisfaction; 100 movie-synopsis narrative environments, each with an initial fact ledger, reference trajectory and commitment set; the evaluated model narrates against an adversarial Gemini-2.5-Flash player, and Gemini-2.5-Flash auditors check every narrator response and end the episode at the first confirmed fact, commitment or player-input conflict, at 100 turns, or when every achievement commitment is met; temperature 0.6, top-p 0.95. Source: arxiv.org. Saturation forecast: Around 2038. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.2 | 32.92 |
| 2 | GPT-4o Mini | 24.8 |
| 3 | Qwen 3 235B A22B | 16.76 |
| 4 | DeepSeek V3.2 | 15.88 |
| 5 | Grok 4.1 Fast | 7.87 |
| 6 | Kimi K2.5 | 2.88 |
Interactive version: theaggregate.ai/benchmark?slug=ncp-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.