NCP-Bench - Satisfied Commitments: leaderboard
Metric: Satisfied commitments (%; share of the narrative commitments in the Satisfied state at termination, averaged over the environments; 100 movie-synopsis narrative environments, each with an initial fact ledger, reference trajectory and commitment set; the evaluated model narrates against an adversarial Gemini-2.5-Flash player, and Gemini-2.5-Flash auditors check every narrator response and end the episode at the first confirmed fact, commitment or player-input conflict, at 100 turns, or when every achievement commitment is met; temperature 0.6, top-p 0.95). Source: arxiv.org. Saturation forecast: Around 2033. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek V3.2 | 13.42 |
| 2 | GPT-5.2 | 11.22 |
| 3 | GPT-4o Mini | 10.9 |
| 4 | Qwen 3 235B A22B | 10.6 |
| 5 | Grok 4.1 Fast | 10.37 |
| 6 | Kimi K2.5 | 5.18 |
Interactive version: theaggregate.ai/benchmark?slug=ncp-bench-satisfied-commitments · How It Works · Data refreshed daily, snapshot 2026-09-29.