PERMA (Noise): leaderboard
Metric: MCQ accuracy (0-1, times 100) on single-domain queries, dialogues with injected in-session noise (vague, fragmented or misleading user turns that keep the intent), PERMA's preference-dependent multiple-choice queries placed along ten simulated users' event-driven interaction timelines (about 34k tokens of dialogue history per user, 580 queries in all); the standalone model reads the full dialogue history with no retrieval or memory compression and must pick the option consistent with the user's evolving persona (options ablate task completion, preference consistency and informational confidence); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 10 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Gemini 2.5 Flash | 87.9 | #237 |
| 2 | Qwen 3 32B | 87.7 | #424 |
| 3 | Kimi K2.5 | 86.5 | #139 |
| 4 | GLM-4.7 Flash | 85.3 | #496 |
| 5 | Llama 3.3 70B Instruct | 82 | #520 |
| 6 | GLM-5 | 81.3 | #137 |
| 7 | MiniMax-M2.5 | 79.7 | #295 |
| 8 | Qwen 2.5 72B Instruct | 79.2 | #436 |
| 9 | GPT-4o Mini | 76.6 | #588 |
| 10 | Qwen2.5-14B-Instruct-1M | 76.6 | #566 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=perma-noise · How It Works · Data refreshed daily, snapshot 2026-10-11.