Qwen2.5-Omni-7B: benchmark results

Alibaba's 7B end-to-end omni-modal model (March 2025) whose Thinker-Talker design takes text, image, audio, and video and streams speech replies. Provider: Alibaba. Released 2025-03-26. Access: Open.

Unified ELO 1545 ± 8, rank #603 of 1605 rated models, from 339 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Afrispeech Semantics - Consistency (AfriSpeech-General)91.31Macro-F1 (%; binary consistent versus inconsistent judgments100
Afrispeech Semantics - Plausibility (AfriSpeech-General)95.69Macro-F1 (%; plausible-but-unsupported versus implausible ju100
Chehre - Distributional Coverage (High Ambiguity, Random Sampling)76Coverage (%; share of the human annotators' rating vectors t100
Chehre - Distributional Coverage (Low Ambiguity, Random Sampling)65Coverage (%; share of the human annotators' rating vectors t100
HELM Audio - Air Bench Foundation74.35EM100
HELM Audio - Vocal Sound90.35PEM100
LLM Stats (FLEURS)95.9Score (%)100
SEA-SpeechBench - Temporal Localization (0-30 s)35.74Span-overlap F1 (%; predict the start and end time at which 100
SEA-SpeechBench - Temporal Localization (30-60 s)19.98Span-overlap F1 (%; predict the start and end time at which 100
SEA-SpeechBench - Temporal Localization (60-120 s)11.32Span-overlap F1 (%; predict the start and end time at which 100
SonicBench - Timbre75Accuracy (%; instrument timbre (NSynth); 100 recognition and100
SonicBench - Loudness75Accuracy (%; loudness (integrated LUFS contrasts of at least98.4

Interactive version: theaggregate.ai/model?slug=qwen2-5-omni-7b · How It Works · Data refreshed daily, snapshot 2026-09-26.