Qwen 3.8 Max (xHigh): benchmark results
Provider: Alibaba. Released 2026-08-03. Access: API.
Unified ELO 1702 ± 1, rank #99 of 1761 rated models, from 16 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| OTIS Mock AIME 2024-25 | 99.44 | Accuracy (%) | 96.9 |
| Epoch AI - Mystery Game Puzzles | 38 | Score | 93.5 |
| FrontierMath - Tiers 1-3 (v2) | 74.74 | Accuracy (%, 285 private v2 problems) | 90.6 |
| Conceptual Reasoning Index - Argument Evaluation (LMCA) | 46.24 | Chance-Corrected Score (0-100) | 87.8 |
| Epoch AI - Dtbench | 92 | Score | 86.9 |
| LLM2014 Logic 2026-08 | 60.05 | Median Score | 86.4 |
| SuperCLUE-Terminal - Overall | 49.49 | Score | 80 |
| Chess Puzzles (Epoch AI) | 29 | Accuracy (%) | 78.5 |
| LLM2014 Logic 2026-09 | 55.61 | Median Score | 75 |
| FrontierMath - Tier 4 (v2) | 46.34 | Accuracy (%, 41 private v2 problems) | 70.6 |
| SimpleQA Verified | 45.8 | Accuracy (%) | 62.3 |
| DeepsecBench | 16.47 | Recall-weighted F2 score (%) | 60 |
Interactive version: theaggregate.ai/model?slug=qwen-3-8-max-xhigh · How It Works · Data refreshed daily, snapshot 2026-09-05.