Qwen2.5-Omni-7B — benchmark results
Alibaba's 7B end-to-end omni-modal model (March 2025) whose Thinker-Talker design takes text, image, audio, and video and streams speech replies. Provider: Alibaba. Released 2025-03-26. Access: Open.
Unified ELO 1482 ± 22, rank #868 of 1776 rated models, from 50 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LLM Stats (FLEURS) | 95.9 | Score (%) | 100 |
| FinBen (Financial LLM) | 33.53 | Average Score | 94.7 |
| LLM Stats (DocVQA) | 95.2 | Score (%) | 90 |
| LLM Stats (TextVQA) | 84.4 | Score (%) | 85.7 |
| KOFFVQA - Korean OCR | 90 | Score (%) | 84 |
| OpenVLM Video - TempCompass | 70.72 | Normalized Score (%) | 83.3 |
| OpenVLM Video | 54.54 | Average Normalized Score (%) | 80.7 |
| OpenVLM Video - MVBench | 69 | Normalized Score (%) | 77.8 |
| KOFFVQA - Document Understanding | 72.67 | Score (%) | 67.9 |
| OpenVLM Video - MLVU | 35.9 | Normalized Score (%) | 66.7 |
| OpenVLM Video - Video-MME (w/o subs) | 64.1 | Normalized Score (%) | 66.7 |
| OpenVLM Video - MMBench-Video | 33 | Normalized Score (%) | 65.9 |
Interactive version: theaggregate.ai/model?slug=qwen2-5-omni-7b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.