ALPBench - Baijiu (Profile F1): leaderboard

Metric: Profile-level F1 (%): the harmonic mean of the shares of predicted and of ground-truth attribute combinations that match exactly (every attribute right), for the next purchased baijiu item's attribute combination (price range, brand, alcohol content and packaging), predicted zero-shot from 100 users' one-year Kuaishou e-commerce interaction histories (clicks, add-to-cart and purchases with titles, selling points and price tiers, up to about 340K tokens); ground truth from verified purchases after a shopping festival cut-off; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 15 models tracked.

Top models

#ModelScoreOverall rank
1Seed 1.820.08#136
2Gemini 2.5 Pro19.2#145
3DeepSeek V3.218.87#198
4Claude Sonnet 4.518.45#138
5Qwen 3 Max (Preview)17.8#224
6GLM-4.517.63#265
7Gemini 2.5 Flash17.53#237
8GLM-4.617.07#246
9MiniMax-M217.02#307
10Qwen 3 235B A22B 2507 Instruct16.9#291
11MiniMax-M2.116.57#283
12DeepSeek R115.83#245
13GPT-5.2 Instant15.37#205
14Qwen 3 30B A3B 2507 (Thinking)13.67#366 (Qwen 3 30B A3B 2507)
15Kimi K26.67#236

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=alpbench-baijiu-profile-f1 · How It Works · Data refreshed daily, snapshot 2026-10-11.