EmoBench-M — leaderboard

Emotional intelligence benchmark for multimodal LLMs evaluating foundational emotion recognition, conversational emotion understanding, and socially complex emotion analysis.

Metric: Average Score. Source: github.com. Status: years away from saturation. 26 models tracked.

Top models

#ModelScore
1Gemini 3 Pro70.5
2GPT-5.266.5
3Gemini 2.0 Flash62.3
4Gemini 1.5 Flash61.3
5Gemini 2.0 Flash (Thinking)60.6
6Qwen 3 VL 8B Instruct60.1
7Qwen 2.5 VL 32B Instruct56.7
8Qwen2.5-Omni-7B53.7
9Qwen2-Audio-7B-Instruct53
10Qwen 2.5 VL 7B Instruct52.9
11InternVL2.5-78B52.4

Interactive version: theaggregate.ai/benchmark?slug=emobench-m · How the rankings work · Data refreshed daily, snapshot 2026-07-22.