AMIGO (Threshold 0.3): leaderboard

Metric: Verified Correct accuracy in percent on AMIGO's dress galleries at similarity threshold 0.3 (770 galleries): the model asks Yes/No attribute questions (at most 20 after upload) to a Qwen3-VL-235B oracle and its final guess is correct and the dialogue leaves the target as the unique best-supported candidate; higher is better. Source: arxiv.org. Saturation forecast: Around December 2027. 5 models tracked.

Top models

#ModelScoreOverall rank
1Gemma 4 31B (Non-reasoning)25.2#195 (Gemma 4 31B)
2Qwen 3.5 397B A17B (Non-reasoning)7.4#141 (Qwen 3.5 397B A17B)
3Gemma 4 12B (Non-reasoning)4.9#529 (Gemma 4 12B)
4Step3 VL 10B1.3#465

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=amigo-threshold-0-3 · How It Works · Data refreshed daily, snapshot 2026-10-11.