Trip+ - Soft Preferences: leaderboard

Metric: Soft-preference satisfaction (0-100), a rate scaled by 100: profile-derived rule scores for pace, comfort, transport, budget and interests, averaged over the plans that pass the response-mode gate, on multi-turn personalized travel planning with traveler profiles, request changes and environment disruptions, the same OpenAI-compatible function-calling scaffold over a fixed travel sandbox for every model, temperature 0 where supported; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 18 models tracked.

Top models

#ModelScore
1Qwen 3.6 27B (Non-reasoning)66.26
2Gemini 3.1 Pro (Preview)63.89
3Kimi K2.663.8
4Qwen 3.5 27B (Non-reasoning)63.75
5DeepSeek V4 Pro62.58
6GPT-5.4 Mini62.07
7Gemini 3 Flash (Preview)61.72
8GLM-4 32B61.41
9GLM-5.161.36
10MiniMax-M2.760.76
11Seed 2.0 Pro60.31
12Hy3-preview59.57
13Gemma 4 26B A4B59.4
14GPT-5.458.54
15DeepSeek V3.258.42

Interactive version: theaggregate.ai/benchmark?slug=trip-plus-soft-preferences · How It Works · Data refreshed daily, snapshot 2026-09-29.