SalesLLM - Chinese (GPT-4o User): leaderboard

Metric: SalesLLM benchmark score (1-10): 0.5 x buying-intent score (end-of-dialogue customer intent from a fine-tuned BERT classifier, five levels mapped to 2, 4, 6, 8, 10) + 0.5 x selling-performance score (0-10 LLM-judge rubric on verbal purchase commitment, concrete next steps, key-information elicitation and objection resolution, credited only on customer-side evidence), on the curated SalesLLM single-product scenarios (Financial Services and Consumer Goods, five difficulty tiers) in Chinese, with GPT-4o as the simulated customer; the model acts as the salesperson for at most 20 rounds at temperature 0.8; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 15 models tracked.

Top models

#ModelScore
1GLM-4.66.74
2DeepSeek V3.1 (Non-reasoning)6.74
3Gemini 3 Pro6.52
4Doubao-1.5-Pro-32k6.5
5Qwen 3 Max6.48
6GPT-4o6.16
7Qwen 2.5 72B6.06
8glm-4-9B6.01
9Gemma 3 27B5.9
10Qwen 3 32B5.81
11Llama 3.3 70B5.74
12Gemini 3 Flash5.71
13MiMo-V2-Flash5.65
14Qwen 3 8B5.4
15GPT-5 Nano5.22

Interactive version: theaggregate.ai/benchmark?slug=salesllm-chinese-gpt-4o-user · How It Works · Data refreshed daily, snapshot 2026-10-07.