Qwen 3.6 27B — benchmark results

Alibaba's dense open Qwen 3.6 27B model, tuned for agentic coding. Provider: Alibaba. Released 2026-04-22. Access: Open.

Unified ELO 1691 ± 12, rank #289 of 1776 rated models, from 125 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
CheXpercept92.2Stage 1 (End-to-End) (self-reported)100
LLM Stats (EmbSpatialBench)84.6Score (%)100
LLM Stats (RefSpatialBench)70Score (%)100
LLM Stats (VideoMME w sub.)87.7Score (%)100
RefSpatialBench70RefSpatialBench (self-reported)100
The Age of Curiosity Meets the Age of AI: Benc4.98Total (self-reported)100
Done, But Not Sure37.8B All (self-reported)94.7
LLM Stats (MVBench)75.5Score (%)93.8
AI for Education Pedagogy - Primary92.96Accuracy (%)93.6
AI for Education Pedagogy - Science92.35Accuracy (%)93.4
AI for Education Pedagogy - Social studies87.27Accuracy (%)93.2
AA-LCR68.7Score (self-reported)92

Interactive version: theaggregate.ai/model?slug=qwen-3-6-27b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.