Qwen 3.5 Plus (2026-04-20): benchmark results
April 20, 2026 Qwen 3.5 Plus snapshot, tracked when sources report the dated API model. Provider: Alibaba. Released 2026-02-16. Access: API.
Unified ELO 1777 ± 48, rank #188 of 2656 rated models, from 14 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| THOR Finding Triage - False Review Load | 23.7 | False positives not suppressed (%, lower is better) | 98.1 |
| TriageBench | 44 | Decision Accuracy (%) | 96.4 |
| SvelteBench | 98.9 | Average pass@1 (%) | 91.3 |
| Wolfram LLM Benchmarking Project | 58.8 | Correct Functionality (%) | 86.4 |
| THOR Finding Triage - CW% | 63.7 | Confidence-weighted classification score (%) | 80.6 |
| Tinybird AI SQL Benchmark - First-Attempt Success Rate | 98 | Questions answered with a valid query on the first attempt ( | 78.1 |
| Tinybird AI SQL Benchmark - Exactness | 51.81 | Result exactness vs human reference queries (0-100) | 73 |
| LM Market Cap LMC Score | 74.8 | LMC Score (0-100) | 72.5 |
| Tinybird AI SQL Benchmark - Success Rate | 100 | Questions answered with a valid query within 3 retries (%) | 72.5 |
| ObviousBench | 91.67 | Answer pass³ (%) | 52.9 |
| THOR Finding Triage - Balanced OTS | 52.8 | Class-balanced operational triage score (%) | 28.7 |
| SnakeBench | 17.2 | TrueSkill Rating | 27.4 |
Interactive version: theaggregate.ai/model?slug=qwen-3-5-plus-2026-04-20 · How It Works · Data refreshed daily, snapshot 2026-09-19.