Qwen 3.6 Plus — benchmark results
Alibaba Qwen 3.6 Plus model for high-end chat, coding, and reasoning workloads. Provider: Alibaba. Released 2026-04-02. Access: API.
Unified ELO 1752 ± 8, rank #193 of 1776 rated models, from 548 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CC-OCR V2 | 75.77 | Average (self-reported) | 100 |
| From Knowing to Doing | 85.29 | Total ret. (self-reported) | 100 |
| JudgeBench Math | 96.43 | Accuracy (%) | 100 |
| LLM Stats (C-Eval) | 93.3 | Score (%) | 100 |
| LLM Stats (CC-OCR) | 83.4 | Score (%) | 100 |
| LLM Stats (DynaMath) | 88 | Score (%) | 100 |
| LLM Stats (MMLongBench-Doc) | 62 | Score (%) | 100 |
| LLM Stats (MMStar) | 83.3 | Score (%) | 100 |
| LLM Stats (ODinW) | 51.8 | Score (%) | 100 |
| LLM Stats (RefCOCO-avg) | 93.5 | Score (%) | 100 |
| RewardBench 2 Math | 91.8 | Accuracy (%) | 99 |
| Evals for Every Language - Language new | 93.33 | Average Score (%) | 98.5 |
Interactive version: theaggregate.ai/model?slug=qwen-3-6-plus · How the rankings work · Data refreshed daily, snapshot 2026-07-22.