ALPBench - Baijiu (Attribute F1): leaderboard

Metric: Attribute-level F1 (%): the F1 of each target attribute computed over all instances, averaged over the category's attributes, for the next purchased baijiu item's attribute combination (price range, brand, alcohol content and packaging), predicted zero-shot from 100 users' one-year Kuaishou e-commerce interaction histories (clicks, add-to-cart and purchases with titles, selling points and price tiers, up to about 340K tokens); ground truth from verified purchases after a shopping festival cut-off; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScoreOverall rank
1Seed 1.865.12#136
2DeepSeek V3.263.13#198
3Claude Sonnet 4.563.11#138
4Gemini 2.5 Pro62.16#145
5Gemini 2.5 Flash59.92#237
6GPT-5.2 Instant59.34#205
7GLM-4.658.55#246
8MiniMax-M256.33#307
9DeepSeek R155.27#245
10MiniMax-M2.154.08#283
11Qwen 3 235B A22B 2507 Instruct53.54#291
12GLM-4.553.35#265
13Qwen 3 Max (Preview)51.46#224
14Qwen 3 30B A3B 2507 (Thinking)47.87#366 (Qwen 3 30B A3B 2507)
15Kimi K237.93#236

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=alpbench-baijiu-attribute-f1 · How It Works · Data refreshed daily, snapshot 2026-10-11.