SalesLLM - English (CustomerLM User): leaderboard
Metric: SalesLLM benchmark score (1-10): 0.5 x buying-intent score (end-of-dialogue customer intent from a fine-tuned BERT classifier, five levels mapped to 2, 4, 6, 8, 10) + 0.5 x selling-performance score (0-10 LLM-judge rubric on verbal purchase commitment, concrete next steps, key-information elicitation and objection resolution, credited only on customer-side evidence), on the curated SalesLLM single-product scenarios (Financial Services and Consumer Goods, five difficulty tiers) in English, with CustomerLM (the authors' Qwen3-8B customer model trained with SFT and DPO on crowdworker sales conversations) as the simulated customer; the model acts as the salesperson for at most 20 rounds at temperature 0.8; higher is better. Source: arxiv.org. Saturation forecast: Around 2032. 15 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash | 6.03 |
| 2 | Gemini 3 Pro | 6.03 |
| 3 | MiMo-V2-Flash | 5.9 |
| 4 | GPT-5 Nano | 5.84 |
| 5 | DeepSeek V3.1 (Non-reasoning) | 5.8 |
| 6 | Qwen 3 8B | 5.79 |
| 7 | Qwen 2.5 72B | 5.63 |
| 8 | Qwen 3 32B | 5.62 |
| 9 | Qwen 3 Max | 5.56 |
| 10 | glm-4-9B | 5.55 |
| 11 | Doubao-1.5-Pro-32k | 5.48 |
| 12 | GLM-4.6 | 5.32 |
| 13 | Llama 3.3 70B | 5.24 |
| 14 | GPT-4o | 5.19 |
| 15 | Gemma 3 27B | 5.09 |
Interactive version: theaggregate.ai/benchmark?slug=salesllm-english-customerlm-user · How It Works · Data refreshed daily, snapshot 2026-10-07.