FindIt - Instance Detection: leaderboard

Metric: Average F1@0.5 (%) over instance detection (find the specific object shown in a reference image) on HR-InsDet easy and hard and RoboTools, 1,000 queries each; each model scored with the box representation, text or JSON output and JSON key that a two-stage probe on Pascal VOC found best for it; boxes Hungarian-matched to the ground truth, a match counts at IoU 0.5 or more (with the right label for multi-label queries); unparseable outputs count as empty; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 10 models tracked.

Top models

#ModelScore
1Qwen 3.5 9B (Thinking)34.8
2Qwen 3.5 9B (Non-reasoning)34.5
3Qwen 3 VL 8B Instruct27.7
4Gemini 2.5 Flash (Non-reasoning)18
5GPT-5.4 (Non-reasoning)16.1
6Qwen 2.5 VL 7B Instruct5.7
7InternVL3-8B0.5
8Gemma 4 E4B0.3
9Claude Sonnet 4.50

Interactive version: theaggregate.ai/benchmark?slug=findit-instance-detection · How It Works · Data refreshed daily, snapshot 2026-09-29.