Qwen2.5-Omni-7B: benchmark results
Alibaba's 7B end-to-end omni-modal model (March 2025) whose Thinker-Talker design takes text, image, audio, and video and streams speech replies. Provider: Alibaba. Released 2025-03-26. Access: Open.
Unified ELO 1545 ± 8, rank #603 of 1605 rated models, from 339 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Afrispeech Semantics - Consistency (AfriSpeech-General) | 91.31 | Macro-F1 (%; binary consistent versus inconsistent judgments | 100 |
| Afrispeech Semantics - Plausibility (AfriSpeech-General) | 95.69 | Macro-F1 (%; plausible-but-unsupported versus implausible ju | 100 |
| Chehre - Distributional Coverage (High Ambiguity, Random Sampling) | 76 | Coverage (%; share of the human annotators' rating vectors t | 100 |
| Chehre - Distributional Coverage (Low Ambiguity, Random Sampling) | 65 | Coverage (%; share of the human annotators' rating vectors t | 100 |
| HELM Audio - Air Bench Foundation | 74.35 | EM | 100 |
| HELM Audio - Vocal Sound | 90.35 | PEM | 100 |
| LLM Stats (FLEURS) | 95.9 | Score (%) | 100 |
| SEA-SpeechBench - Temporal Localization (0-30 s) | 35.74 | Span-overlap F1 (%; predict the start and end time at which | 100 |
| SEA-SpeechBench - Temporal Localization (30-60 s) | 19.98 | Span-overlap F1 (%; predict the start and end time at which | 100 |
| SEA-SpeechBench - Temporal Localization (60-120 s) | 11.32 | Span-overlap F1 (%; predict the start and end time at which | 100 |
| SonicBench - Timbre | 75 | Accuracy (%; instrument timbre (NSynth); 100 recognition and | 100 |
| SonicBench - Loudness | 75 | Accuracy (%; loudness (integrated LUFS contrasts of at least | 98.4 |
Interactive version: theaggregate.ai/model?slug=qwen2-5-omni-7b · How It Works · Data refreshed daily, snapshot 2026-09-26.