COIN-Bench - Depth: leaderboard

Metric: Depth score (%): share of COIN-Tree node weight the questionnaire covers at each of five levels (usage scenarios, aspects, feelings, comparisons, tendencies), averaged over the levels, for questionnaires the model writes from each product's consumer discussions, matched by sentence embeddings to COIN-Tree, a weighted intent tree that GPT-4o extracted from the same discussions; higher is better. Source: arxiv.org. Saturation forecast: Around 2034. 22 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.211.61#105
2Gemini 3 Flash11.54#93
3Qwen 3 30B A3B11.33#488
4GPT-4.111.01#240
5GPT-510.78#91
6Qwen 3 32B10.52#424
7GPT-4o10.44#333
8DeepSeek R1 Distill Qwen 14B10.4#828
9Qwen 2.5 32B Instruct10.21#491
10Claude 3.5 Sonnet10.16#337
11DeepSeek R1 Distill Qwen 32B10.07#640
12Qwen 2.5 72B Instruct10#436
13O39.51#121
14Qwen 3 8B9.13#667
15Qwen 2.5 14B Instruct8.83#634

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=coin-bench-depth · How It Works · Data refreshed daily, snapshot 2026-10-11.