EPIC-Bench - Visual Matching: leaderboard

Metric: Navigation score (0-100) on visual matching (localizing an object marked in one view in another view of the same scene), on EPIC-Bench (6,661 human-annotated image, text and mask tuples from 25 public datasets across 23 fine-grained embodied perception tasks; the model outputs bounding boxes, counts, path points or feasibility judgements that are scored against the masks), zero-shot, averaged over six runs for local models and two for API models; higher is better. Source: arxiv.org. Saturation forecast: Around December 2027. 89 models tracked.

Top models

#ModelScore
1Gemini 3 Pro60.08
2Gemini 3.1 Pro (Preview)59.06
3GPT-5.5 (Non-reasoning)51.65
4Gemini 3 Flash (Preview)49.92
5Qwen 3 VL 235B A22B (Thinking)49.01
6Gemini 3 Flash (Preview) (Non-reasoning)47.25
7Qwen 3.5 27B46.77
8GPT-5 Mini45.97
9Seed 1.845.9
10O4 Mini45.63
11Qwen 3.5 397B A17B45.35
12Qwen 3.5 Plus45.21
13Qwen 3.5 122B A10B45.19
14GPT-5.1 (Non-reasoning)44.69
15GPT-5.4 (Non-reasoning)44.35

Interactive version: theaggregate.ai/benchmark?slug=epic-bench-visual-matching · How It Works · Data refreshed daily, snapshot 2026-10-07.