Qwen 3.8 27B: benchmark results
Provider: Alibaba. Released 2026-08-14. Access: Open.
Unified ELO 1689 ± 1, rank #70 of 1392 rated models, from 137 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LiveBench Python | 75 | Score | 98.1 |
| AI for Education Pedagogy - Science | 93.44 | Accuracy (%) | 96 |
| OpenRouter Tau2-Bench Airline | 78.7 | Accuracy (%) | 95.8 |
| OSWorld-Verified | 84.3 | Success rate (self-reported) | 95.4 |
| Bullshit Benchmark | 76.4 | BS Detection Rate (%) | 93.6 |
| LLM Stats Score | 45.17 | LLM Stats Score (conservative rating) | 91.7 |
| RealWorldQA | 85.9 | RealWorldQA (self-reported) | 91.7 |
| BenchmarkList ECI | 143.19 | Capability Index (ECI) | 91.5 |
| WebDev Arena | 1594.78 | Arena Score | 91.5 |
| LLM Stats (MathVision) | 94.6 | Score (%) | 91.2 |
| LLM Stats (CharXiv-R) | 90.2 | Score (%) | 90.7 |
| WebDev Arena (Reference-Based Design) | 1614 | Arena Score | 89.3 |
Interactive version: theaggregate.ai/model?slug=qwen-3-8-27b · How It Works · Data refreshed daily, snapshot 2026-09-05.