MMShopBench (No Verification) - Judge@3: leaderboard

Metric: Judge success at 3 (%; share of cases in which any of the first three recommended products satisfies the annotated purchase intent and every mandatory requirement according to a GPT-5.5 multimodal judge given the product evidence; multimodal multi-turn shopping requests from real assistant logs, answered by a tool-using agent over a frozen offline catalog with text and region-aware image retrieval (at most eight tool steps), without the evidence-grounded verification step). Source: arxiv.org. Saturation forecast: Around December 2026. 7 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)69.2
2Claude Opus 4.8 (Thinking)65.4
3Kimi K2.6 (Thinking)43.6
4MiniMax-M2.724.9
5Qwen 3.5 9B (Non-reasoning)11.1
6Qwen 3.5 27B (Non-reasoning)10
7Qwen 3.5 122B A10B (Non-reasoning)5.2

Interactive version: theaggregate.ai/benchmark?slug=mmshopbench-no-verification-judge-3 · How It Works · Data refreshed daily, snapshot 2026-09-29.