Qwen3.8 Omni Flash: benchmark results
Provider: Alibaba. Access: Open.
Unified ELO 1803 ± 34, rank #60 of 1632 rated models, from 17 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| IFBench | 81.5 | Score (%) | 100 |
| LLM Stats (MMAU) | 81.8 | Score (%) | 100 |
| LLM Stats (WildClawBench) | 71 | Score (%) | 100 |
| RealWorldQA | 87.7 | RealWorldQA (self-reported) | 95.9 |
| LVBench | 76.9 | Score (self-reported) | 82.9 |
| LLM Stats Score | 37.84 | LLM Stats Score (conservative rating) | 79.4 |
| THOR Finding Triage - Critical Miss Rate | 0 | True positives suppressed as false positives (%, lower is be | 77.4 |
| Tinybird AI SQL Benchmark - Exactness | 52.72 | Result exactness vs human reference queries (0-100) | 75.6 |
| ERQA | 71 | Score (%) | 75 |
| NL2Repo | 48.9 | Score (self-reported) | 57 |
| THOR Finding Triage - CW% | 56.4 | Confidence-weighted classification score (%) | 50 |
| THOR Finding Triage - Balanced OTS | 56.9 | Class-balanced operational triage score (%) | 45.8 |
Interactive version: theaggregate.ai/model?slug=qwen3-8-omni-flash · How It Works · Data refreshed daily, snapshot 2026-10-08.