Qwen 3.5 Plus (2026-02-15): benchmark results
Provider: Alibaba. Released 2026-02-16. Access: API.
Unified ELO 1712 ± 25, rank #217 of 1605 rated models, from 26 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EgoArgus (Assistance Decision) - Contradictory | 91.01 | Assistance-decision F1 (%; whether help is required, from th | 100 |
| FireWorldBench - Real-World-Aligned (Structured Records) | 55.35 | Completion accuracy (%; equal-weight mean of P1-P5 on the 71 | 92.9 |
| LiveSecBench | 84.34 | Overall Score (%) | 92.9 |
| ObviousBench | 98.61 | Answer pass³ (%) | 86.9 |
| FireWorldBench - Counterfactual Intervention Reasoning (Structured Records) | 44.92 | Completion accuracy (%; P5, counterfactual intervention reas | 85.7 |
| FireWorldBench - Temporal Evolution Forecasting (Structured Records) | 60.77 | Completion accuracy (%; P1, temporal evolution forecasting i | 85.7 |
| EgoArgus (Assistance Decision) - Multimodal Grounded | 93.4 | Assistance-decision F1 (%; whether help is required, from th | 83.3 |
| FireWorldBench - Counterfactual Intervention Reasoning (Rendered Images) | 37.83 | Completion accuracy (%; P5, counterfactual intervention reas | 77.8 |
| FireWorldBench - Physical Field Perception and Grounding (Rendered Images) | 40.77 | Completion accuracy (%; P2, physical field perception and gr | 77.8 |
| Wolfram LLM Benchmarking Project | 53.1 | Correct Functionality (%) | 75.6 |
| SnakeBench | 26.5 | TrueSkill Rating | 75.1 |
| FireWorldBench (Structured Records) | 47.25 | Completion accuracy (%; equal-weight mean of the five physic | 71.4 |
Interactive version: theaggregate.ai/model?slug=qwen-3-5-plus-2026-02-15 · How It Works · Data refreshed daily, snapshot 2026-09-26.