COIN-Bench - Depth: leaderboard
Metric: Depth score (%): share of COIN-Tree node weight the questionnaire covers at each of five levels (usage scenarios, aspects, feelings, comparisons, tendencies), averaged over the levels, for questionnaires the model writes from each product's consumer discussions, matched by sentence embeddings to COIN-Tree, a weighted intent tree that GPT-4o extracted from the same discussions; higher is better. Source: arxiv.org. Saturation forecast: Around 2034. 22 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GPT-5.2 | 11.61 | #105 |
| 2 | Gemini 3 Flash | 11.54 | #93 |
| 3 | Qwen 3 30B A3B | 11.33 | #488 |
| 4 | GPT-4.1 | 11.01 | #240 |
| 5 | GPT-5 | 10.78 | #91 |
| 6 | Qwen 3 32B | 10.52 | #424 |
| 7 | GPT-4o | 10.44 | #333 |
| 8 | DeepSeek R1 Distill Qwen 14B | 10.4 | #828 |
| 9 | Qwen 2.5 32B Instruct | 10.21 | #491 |
| 10 | Claude 3.5 Sonnet | 10.16 | #337 |
| 11 | DeepSeek R1 Distill Qwen 32B | 10.07 | #640 |
| 12 | Qwen 2.5 72B Instruct | 10 | #436 |
| 13 | O3 | 9.51 | #121 |
| 14 | Qwen 3 8B | 9.13 | #667 |
| 15 | Qwen 2.5 14B Instruct | 8.83 | #634 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=coin-bench-depth · How It Works · Data refreshed daily, snapshot 2026-10-11.