EPIC-Bench - Feasible Path: leaderboard

Metric: Navigation path score (0-100) on feasible path planning (a point path from the camera or a marked region to the target that stays on traversable ground), on EPIC-Bench (6,661 human-annotated image, text and mask tuples from 25 public datasets across 23 fine-grained embodied perception tasks; the model outputs bounding boxes, counts, path points or feasibility judgements that are scored against the masks), zero-shot, averaged over six runs for local models and two for API models; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 89 models tracked.

Top models

#ModelScore
1Gemini 3 Pro70.82
2Gemini 3.1 Pro (Preview)70.32
3Claude Sonnet 4.6 (Thinking)68.46
4Qwen 3.5 122B A10B63.47
5Qwen 3.6 Plus62.12
6Qwen 3.5 35B A3B61.17
7Qwen 3.5 397B A17B60.11
8Qwen 3.6 27B (Non-reasoning)59.15
9Qwen 3.5 Plus57.89
10Qwen 3.5 27B (Non-reasoning)53.5
11Seed 1.853.43
12Gemini 3 Flash (Preview)52.63
13Gemini 2.5 Flash (Non-reasoning)51.98
14Qwen 3 VL 235B A22B (Thinking)51.19
15GLM-4.6V51.11

Interactive version: theaggregate.ai/benchmark?slug=epic-bench-feasible-path · How It Works · Data refreshed daily, snapshot 2026-10-07.