Qwen 3.5 Flash (Thinking): benchmark results
Provider: Alibaba. Released 2026-02-16. Access: API.
Unified ELO 1636 ± 20, rank #558 of 2088 rated models, from 23 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| MT-JailBench - Crescendo | 28.3 | Attack success rate (%) of the Crescendo multi-turn jailbrea | 90 |
| MT-JailBench - XTeaming | 23.27 | Attack success rate (%) of the XTeaming multi-turn jailbreak | 90 |
| OccuBench - Healthcare & Life Sciences | 76 | Completion rate (%) on the Healthcare & Life Sciences indust | 78.6 |
| OccuBench - Science & Research | 69 | Completion rate (%) on the Science & Research industry's tas | 71.4 |
| OccuBench - Commerce & Consumer | 67 | Completion rate (%) on the Commerce & Consumer industry's ta | 67.9 |
| LLM2014 Logic 2026-03 | 36.36 | Median Score | 53.7 |
| LLM2014 Logic 2026-02 | 36.59 | Median Score | 51.1 |
| OccuBench - Agriculture & Environment | 61 | Completion rate (%) on the Agriculture & Environment industr | 42.9 |
| GroupTravelBench - Easy | 6.2 | Group Utility (GU, unnormalized points per user, unbounded a | 37.5 |
| GroupTravelBench - Hard | 7.65 | Group Utility (GU, unnormalized points per user, unbounded a | 37.5 |
| GroupTravelBench - LLM Judge | 43.7 | LLM-judge process score (0-100): mean of five 1-5 ratings (h | 37.5 |
| GroupTravelBench - Medium | 7.05 | Group Utility (GU, unnormalized points per user, unbounded a | 37.5 |
Interactive version: theaggregate.ai/model?slug=qwen-3-5-flash-thinking · How It Works · Data refreshed daily, snapshot 2026-10-07.