CUE-Bench - Pragmatic Intent: leaderboard

Metric: Accuracy (%; pragmatic intent classification; over CUE-Bench's Chinese-discourse instances with context; zero-shot direct prompting). Source: arxiv.org. Saturation forecast: Around 2032. 6 models tracked.

Top models

#ModelScore
1DeepSeek V4 Flash32.5
2GLM-5.131.2
3GPT-4o Mini28.1
4Llama 4 Maverick27.3
5Llama 3.1 8B23.2
6Qwen 3 8B19.6

Interactive version: theaggregate.ai/benchmark?slug=cue-bench-pragmatic-intent · How It Works · Data refreshed daily, snapshot 2026-09-26.