MMShopBench (No Verification) - Identifier Match@3: leaderboard
Metric: Identifier match at 3 (%; share of cases in which any of the first three recommended products is one of the human-verified target product identifiers; multimodal multi-turn shopping requests from real assistant logs, answered by a tool-using agent over a frozen offline catalog with text and region-aware image retrieval (at most eight tool steps), without the evidence-grounded verification step). Source: arxiv.org. Saturation forecast: Around December 2026. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 61.6 |
| 2 | Claude Opus 4.8 (Thinking) | 57.4 |
| 3 | Kimi K2.6 (Thinking) | 35.6 |
| 4 | MiniMax-M2.7 | 21.5 |
| 5 | Qwen 3.5 27B (Non-reasoning) | 10 |
| 6 | Qwen 3.5 9B (Non-reasoning) | 9.7 |
| 7 | Qwen 3.5 122B A10B (Non-reasoning) | 4.8 |
Interactive version: theaggregate.ai/benchmark?slug=mmshopbench-no-verification-identifier-match-3 · How It Works · Data refreshed daily, snapshot 2026-09-29.