ERQA: leaderboard

Embodied Reasoning Question Answering benchmark with visual questions about space, trajectories, actions, state estimation, and multi-view reasoning.

Metric: Score (%). Source: llm-stats.com. Status: years away from saturation. 26 models tracked.

Top models

#ModelScore
1Qwen 3.8 Max77.8
2Qwen3.8-Flash-Next72.3
3Qwen 3.8 Flash72.3
4Seed 2.1 Pro72
5Seed 2.1 Turbo71.3
6Qwen 3.7 Plus69.8
7GPT-565.7
8Qwen 3.6 Plus65.7
9Qwen 3.8 27B65.5
10Qwen 3.5 35B A3B64.8
11Muse Spark64.7
12O364
13Qwen 3.6 27B62.5
14Qwen 3.5 122B A10B62
15Qwen 3.5 27B60.5

Interactive version: theaggregate.ai/benchmark?slug=erqa · How It Works · Data refreshed daily, snapshot 2026-09-05.