YapBench — leaderboard

Measures LLM verbosity on brevity-ideal prompts. Models are scored on excess length beyond a minimal baseline answer using YapScore and YapIndex metrics.

Metric: YapIndex (lower is better). Source: huggingface.co. 108 models tracked.

Top models

#ModelScore
1GLM-4.51427
2nova-2-lite-v1896.3
3Qwen 3 32B730.2
4Qwen 3 VL 235B A22B (Thinking)666.3
5OLMo 3.1 32B (Thinking)563
6Qwen 3 14B558
7Gemini 2.5 Pro543.8
8Qwen 3 VL 8B Instruct518
9Claude Sonnet 5507.7
10Gemini 2.5 Flash Lite476.2
11DeepSeek V3.2 (Thinking)462.3
12Devstral Small 2449
13Mistral Large 3444
14GLM-4.6V443.8
15DeepSeek V3.2426.2

Interactive version: theaggregate.ai/benchmark?slug=yapbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.