Qwen 3.6 Plus (Thinking): benchmark results

Provider: Alibaba. Released 2026-04-02. Access: API.

Unified ELO 1667 ± 1, rank #172 of 1935 rated models, from 22 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HPR-Bench - Tag Inference73.29Avg@10 Tag Accuracy (%; mean of 10 samples over 18 profile t100
HPR-Bench - Tag Inference - Geographic Context79.93Avg@10 Tag Accuracy (%; mean of 10 samples)100
HPR-Bench - Tag Inference - Life Stage79.07Avg@10 Tag Accuracy (%; mean of 10 samples)100
HPR-Bench - Tag Inference - Household Context59.65Avg@10 Tag Accuracy (%; mean of 10 samples)92.9
SuperCLUE-VLM (April 2026) - Overall86.73Score81.2
MMCL-Bench20.6Overall (self-reported)80
HPR-Bench - Tag Inference - Lifestyle Indicators69.91Avg@10 Tag Accuracy (%; mean of 10 samples)78.6
Korean CSAT 2026 (Easy Mode) - Chemistry I47Points (out of 50)74.1
Korean CSAT 2026 (Easy Mode) - Physics I39Points (out of 50)71.1
LLM2014 Logic 2026-0455.35Median Score70
Korean CSAT 2026 (Easy Mode) - English97Points (out of 100)66
Korean CSAT 2026 (Easy Mode) - Society and Culture41Points (out of 50)64.3

Interactive version: theaggregate.ai/model?slug=qwen-3-6-plus-thinking · How It Works · Data refreshed daily, snapshot 2026-09-25.