SalesLLM - Multi-Product Chinese (GPT-4o User): leaderboard
Metric: SalesLLM benchmark score (1-10): 0.5 x buying-intent score (end-of-dialogue customer intent from a fine-tuned BERT classifier, five levels mapped to 2, 4, 6, 8, 10) + 0.5 x selling-performance score (0-10 LLM-judge rubric on verbal purchase commitment, concrete next steps, key-information elicitation and objection resolution, credited only on customer-side evidence), on the multi-product SalesLLM scenarios (six persona-matched products per scenario) in Chinese, with GPT-4o as the simulated customer; the model acts as the salesperson for at most 20 rounds at temperature 0.8; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Doubao-1.5-Pro-32k | 7.46 |
| 2 | Gemini 3 Pro | 7.17 |
| 3 | DeepSeek V3.1 (Non-reasoning) | 6.47 |
Interactive version: theaggregate.ai/benchmark?slug=salesllm-multi-product-chinese-gpt-4o-user · How It Works · Data refreshed daily, snapshot 2026-10-07.