ALPBench - Apparel (Attribute F1): leaderboard

Metric: Attribute-level F1 (%): the F1 of each target attribute computed over all instances, averaged over the category's attributes, for the next purchased apparel item's attribute combination (price range, fit, thickness, category and sleeve length), predicted zero-shot from 100 users' one-year Kuaishou e-commerce interaction histories (clicks, add-to-cart and purchases with titles, selling points and price tiers, up to about 340K tokens); ground truth from verified purchases after a shopping festival cut-off; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro42.66#145
2Seed 1.840.42#136
3GPT-5.2 Instant39.76#205
4Qwen 3 Max (Preview)38.89#224
5MiniMax-M2.137.55#283
6MiniMax-M234.5#307
7Gemini 2.5 Flash34.4#237
8GLM-4.529.88#265
9GLM-4.629.84#246
10Claude Sonnet 4.529.33#138
11Qwen 3 235B A22B 2507 Instruct23.98#291
12DeepSeek V3.223.2#198
13DeepSeek R119.61#245
14Qwen 3 30B A3B 2507 (Thinking)19.38#366 (Qwen 3 30B A3B 2507)
15Kimi K212.74#236

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=alpbench-apparel-attribute-f1 · How It Works · Data refreshed daily, snapshot 2026-10-11.