EPIC-Bench - Embodied Compositional Attributes: leaderboard

Metric: Target localization score (0-100) on the embodied compositional tasks (targets described by combined attributes, part-whole relations and interaction context), on EPIC-Bench (6,661 human-annotated image, text and mask tuples from 25 public datasets across 23 fine-grained embodied perception tasks; the model outputs bounding boxes, counts, path points or feasibility judgements that are scored against the masks), zero-shot, averaged over six runs for local models and two for API models; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 89 models tracked.

Top models

#ModelScore
1Gemini 3 Pro53.9
2Gemini 3.1 Pro (Preview)53.04
3Qwen 3 VL 235B A22B (Thinking)48.12
4Seed 1.846.1
5GPT-5.5 (Non-reasoning)45.19
6Qwen 3.5 397B A17B (Non-reasoning)43.45
7Qwen 3.5 122B A10B (Non-reasoning)42.31
8Qwen 3.5 27B (Non-reasoning)42.23
9Qwen 3.5 397B A17B42.1
10Qwen 3.5 Plus41.82
11Qwen 3.6 27B (Non-reasoning)41.46
12Qwen 3.5 122B A10B41.16
13Qwen 3.6 Plus41.07
14Qwen 2.5 VL 72B Instruct40.92
15Qwen 3 VL 30B A3B (Thinking)40.87

Interactive version: theaggregate.ai/benchmark?slug=epic-bench-embodied-compositional-attributes · How It Works · Data refreshed daily, snapshot 2026-10-07.