Qwen 3.5 Plus (2026-02-15): benchmark results

Provider: Alibaba. Released 2026-02-16. Access: API.

Unified ELO 1712 ± 25, rank #217 of 1605 rated models, from 26 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EgoArgus (Assistance Decision) - Contradictory91.01Assistance-decision F1 (%; whether help is required, from th100
FireWorldBench - Real-World-Aligned (Structured Records)55.35Completion accuracy (%; equal-weight mean of P1-P5 on the 7192.9
LiveSecBench84.34Overall Score (%)92.9
ObviousBench98.61Answer pass³ (%)86.9
FireWorldBench - Counterfactual Intervention Reasoning (Structured Records)44.92Completion accuracy (%; P5, counterfactual intervention reas85.7
FireWorldBench - Temporal Evolution Forecasting (Structured Records)60.77Completion accuracy (%; P1, temporal evolution forecasting i85.7
EgoArgus (Assistance Decision) - Multimodal Grounded93.4Assistance-decision F1 (%; whether help is required, from th83.3
FireWorldBench - Counterfactual Intervention Reasoning (Rendered Images)37.83Completion accuracy (%; P5, counterfactual intervention reas77.8
FireWorldBench - Physical Field Perception and Grounding (Rendered Images)40.77Completion accuracy (%; P2, physical field perception and gr77.8
Wolfram LLM Benchmarking Project53.1Correct Functionality (%)75.6
SnakeBench26.5TrueSkill Rating75.1
FireWorldBench (Structured Records)47.25Completion accuracy (%; equal-weight mean of the five physic71.4

Interactive version: theaggregate.ai/model?slug=qwen-3-5-plus-2026-02-15 · How It Works · Data refreshed daily, snapshot 2026-09-26.