Qwen 3.5 Plus (Non-reasoning): benchmark results
Provider: Alibaba. Released 2026-02-16. Access: API.
Unified ELO 1649 ± 23, rank #516 of 2055 rated models, from 21 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| UAV-DualCog - Flight Behavior Recognition (Atomic) | 31.7 | Atomic-level behavior accuracy (%; the atomic flight actions | 76.9 |
| RuleWeaver - Cross-Source - Rule Recall | 55.83 | Gold-rule recall (%) | 70 |
| RuleWeaver - Same-Source - Rule Precision | 66.32 | Correctly applied share of cited rules (%) | 70 |
| UAV-DualCog - Landmark-Driven Action Decision | 48.9 | Answer accuracy (%; 1,024 questions asking which way the UAV | 65.7 |
| UAV-DualCog - Landmark Visibility Counting | 48.3 | Visibility-count accuracy (%; 1,024 flight videos in which t | 60.7 |
| RuleWeaver - Cross-Source - Rubric Score | 34.15 | Judge rubric score (0-100) | 60 |
| RuleWeaver - Same-Source - Rubric Score | 37.75 | Judge rubric score (0-100) | 60 |
| RuleWeaver - Same-Source - Rule Recall | 66.67 | Gold-rule recall (%) | 60 |
| UAV-DualCog - Future Observation Prediction | 29 | Answer accuracy (%; 1,024 questions asking which view the UA | 57.4 |
| UAV-DualCog - Landmark-Relative Direction | 47.6 | Answer accuracy (%; 1,024 questions asking where a landmark | 55.7 |
| Epoch AI - Mystery Game Puzzles | 17 | Score | 46.9 |
| RuleWeaver - Cross-Source - Rule Precision | 64.27 | Correctly applied share of cited rules (%) | 40 |
Interactive version: theaggregate.ai/model?slug=qwen-3-5-plus-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-29.