Qwen 3.5 Plus (Non-reasoning): benchmark results

Provider: Alibaba. Released 2026-02-16. Access: API.

Unified ELO 1649 ± 23, rank #516 of 2055 rated models, from 21 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
UAV-DualCog - Flight Behavior Recognition (Atomic)31.7Atomic-level behavior accuracy (%; the atomic flight actions76.9
RuleWeaver - Cross-Source - Rule Recall55.83Gold-rule recall (%)70
RuleWeaver - Same-Source - Rule Precision66.32Correctly applied share of cited rules (%)70
UAV-DualCog - Landmark-Driven Action Decision48.9Answer accuracy (%; 1,024 questions asking which way the UAV65.7
UAV-DualCog - Landmark Visibility Counting48.3Visibility-count accuracy (%; 1,024 flight videos in which t60.7
RuleWeaver - Cross-Source - Rubric Score34.15Judge rubric score (0-100)60
RuleWeaver - Same-Source - Rubric Score37.75Judge rubric score (0-100)60
RuleWeaver - Same-Source - Rule Recall66.67Gold-rule recall (%)60
UAV-DualCog - Future Observation Prediction29Answer accuracy (%; 1,024 questions asking which view the UA57.4
UAV-DualCog - Landmark-Relative Direction47.6Answer accuracy (%; 1,024 questions asking where a landmark 55.7
Epoch AI - Mystery Game Puzzles17Score46.9
RuleWeaver - Cross-Source - Rule Precision64.27Correctly applied share of cited rules (%)40

Interactive version: theaggregate.ai/model?slug=qwen-3-5-plus-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-29.