Qwen 3.6 Plus (Thinking): benchmark results
Provider: Alibaba. Released 2026-04-02. Access: API.
Unified ELO 1667 ± 1, rank #172 of 1935 rated models, from 22 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HPR-Bench - Tag Inference | 73.29 | Avg@10 Tag Accuracy (%; mean of 10 samples over 18 profile t | 100 |
| HPR-Bench - Tag Inference - Geographic Context | 79.93 | Avg@10 Tag Accuracy (%; mean of 10 samples) | 100 |
| HPR-Bench - Tag Inference - Life Stage | 79.07 | Avg@10 Tag Accuracy (%; mean of 10 samples) | 100 |
| HPR-Bench - Tag Inference - Household Context | 59.65 | Avg@10 Tag Accuracy (%; mean of 10 samples) | 92.9 |
| SuperCLUE-VLM (April 2026) - Overall | 86.73 | Score | 81.2 |
| MMCL-Bench | 20.6 | Overall (self-reported) | 80 |
| HPR-Bench - Tag Inference - Lifestyle Indicators | 69.91 | Avg@10 Tag Accuracy (%; mean of 10 samples) | 78.6 |
| Korean CSAT 2026 (Easy Mode) - Chemistry I | 47 | Points (out of 50) | 74.1 |
| Korean CSAT 2026 (Easy Mode) - Physics I | 39 | Points (out of 50) | 71.1 |
| LLM2014 Logic 2026-04 | 55.35 | Median Score | 70 |
| Korean CSAT 2026 (Easy Mode) - English | 97 | Points (out of 100) | 66 |
| Korean CSAT 2026 (Easy Mode) - Society and Culture | 41 | Points (out of 50) | 64.3 |
Interactive version: theaggregate.ai/model?slug=qwen-3-6-plus-thinking · How It Works · Data refreshed daily, snapshot 2026-09-25.