Qwen 3.5 Flash (Thinking): benchmark results

Provider: Alibaba. Released 2026-02-16. Access: API.

Unified ELO 1636 ± 20, rank #558 of 2088 rated models, from 23 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
MT-JailBench - Crescendo28.3Attack success rate (%) of the Crescendo multi-turn jailbrea90
MT-JailBench - XTeaming23.27Attack success rate (%) of the XTeaming multi-turn jailbreak90
OccuBench - Healthcare & Life Sciences76Completion rate (%) on the Healthcare & Life Sciences indust78.6
OccuBench - Science & Research69Completion rate (%) on the Science & Research industry's tas71.4
OccuBench - Commerce & Consumer67Completion rate (%) on the Commerce & Consumer industry's ta67.9
LLM2014 Logic 2026-0336.36Median Score53.7
LLM2014 Logic 2026-0236.59Median Score51.1
OccuBench - Agriculture & Environment61Completion rate (%) on the Agriculture & Environment industr42.9
GroupTravelBench - Easy6.2Group Utility (GU, unnormalized points per user, unbounded a37.5
GroupTravelBench - Hard7.65Group Utility (GU, unnormalized points per user, unbounded a37.5
GroupTravelBench - LLM Judge43.7LLM-judge process score (0-100): mean of five 1-5 ratings (h37.5
GroupTravelBench - Medium7.05Group Utility (GU, unnormalized points per user, unbounded a37.5

Interactive version: theaggregate.ai/model?slug=qwen-3-5-flash-thinking · How It Works · Data refreshed daily, snapshot 2026-10-07.