Qwen 3.8 Flash (Max): benchmark results
Provider: Alibaba. Released 2026-08-26. Access: API.
Unified ELO 1716 ± 1, rank #97 of 3078 rated models, from 12 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SuperCLUE General (July 2026) - Agentic Task Planning | 88.61 | Score | 82.4 |
| SuperCLUE General (July 2026) - Hallucination Control | 86.08 | Score | 79.4 |
| Bug Hunt Bench - VS Code Extension | 13 | Planted Bugs Fixed (out of 45) | 75.4 |
| AI Coding Daily (OpenCode) - React-TS Code Quality | 17.67 | React-TS Code Quality (max 20) points, LLM-judged rubric sco | 71.1 |
| Bug Hunt Bench | 26 | Planted Bugs Fixed (out of 105) | 69.4 |
| AI Coding Daily (OpenCode) - Total | 47.22 | Total points (max 60) | 63.2 |
| Bug Hunt Bench - LMS | 13 | Planted Bugs Fixed (out of 60) | 62.7 |
| SuperCLUE General (July 2026) - Math Reasoning | 75.44 | Score | 61.8 |
| SuperCLUE General (July 2026) - Overall | 66.41 | Score | 58.8 |
| AI Coding Daily (OpenCode) - Laravel Code Quality | 16.35 | Laravel Code Quality (max 20) points, LLM-judged rubric scor | 31.6 |
| SuperCLUE General (July 2026) - Precise Instruction Following | 30.48 | Score | 29.4 |
| SuperCLUE General (July 2026) - Science Reasoning | 68.42 | Score | 23.5 |
Interactive version: theaggregate.ai/model?slug=qwen-3-8-flash-max · How It Works · Data refreshed daily, snapshot 2026-09-19.