Qwen2.5-Omni-7B — benchmark results

Alibaba's 7B end-to-end omni-modal model (March 2025) whose Thinker-Talker design takes text, image, audio, and video and streams speech replies. Provider: Alibaba. Released 2025-03-26. Access: Open.

Unified ELO 1482 ± 22, rank #868 of 1776 rated models, from 50 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LLM Stats (FLEURS)95.9Score (%)100
FinBen (Financial LLM)33.53Average Score94.7
LLM Stats (DocVQA)95.2Score (%)90
LLM Stats (TextVQA)84.4Score (%)85.7
KOFFVQA - Korean OCR90Score (%)84
OpenVLM Video - TempCompass70.72Normalized Score (%)83.3
OpenVLM Video54.54Average Normalized Score (%)80.7
OpenVLM Video - MVBench69Normalized Score (%)77.8
KOFFVQA - Document Understanding72.67Score (%)67.9
OpenVLM Video - MLVU35.9Normalized Score (%)66.7
OpenVLM Video - Video-MME (w/o subs)64.1Normalized Score (%)66.7
OpenVLM Video - MMBench-Video33Normalized Score (%)65.9

Interactive version: theaggregate.ai/model?slug=qwen2-5-omni-7b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.