B2V-Bench (With Few-Shot Examples): leaderboard

Metric: Macro-F1 (%; multi-label classification of the consumer values, 28 labels of the E-commerce Consumption Value Taxonomy plus none, expressed in each of 544 Taobao purchase-decision episodes labelled from scratch by psychology experts; the model reads the GPT-5.4-written summary of the pre-purchase behaviour log; prompt adds the behavioural value descriptions and few-shot examples). Source: arxiv.org. Saturation forecast: Around 2028. 8 models tracked.

Top models

#ModelScore
1GPT-541.7
2GLM-535.8
3DeepSeek V3.235
4Kimi K2.630.4
5Qwen 3 235B A22B28.8
6Claude Haiku 4.527.5
7GPT-4o Mini22

Interactive version: theaggregate.ai/benchmark?slug=b2v-bench-with-few-shot-examples · How It Works · Data refreshed daily, snapshot 2026-09-26.