Qwen3.8 Omni Flash: benchmark results

Provider: Alibaba. Access: Open.

Unified ELO 1803 ± 34, rank #60 of 1632 rated models, from 17 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
IFBench81.5Score (%)100
LLM Stats (MMAU)81.8Score (%)100
LLM Stats (WildClawBench)71Score (%)100
RealWorldQA87.7RealWorldQA (self-reported)95.9
LVBench76.9Score (self-reported)82.9
LLM Stats Score37.84LLM Stats Score (conservative rating)79.4
THOR Finding Triage - Critical Miss Rate0True positives suppressed as false positives (%, lower is be77.4
Tinybird AI SQL Benchmark - Exactness52.72Result exactness vs human reference queries (0-100)75.6
ERQA71Score (%)75
NL2Repo48.9Score (self-reported)57
THOR Finding Triage - CW%56.4Confidence-weighted classification score (%)50
THOR Finding Triage - Balanced OTS56.9Class-balanced operational triage score (%)45.8

Interactive version: theaggregate.ai/model?slug=qwen3-8-omni-flash · How It Works · Data refreshed daily, snapshot 2026-10-08.