TravelEval - Preference-Weighted Experience: leaderboard

Metric: Profit: mean over attractions of objective score times the preference-match weight (paper units, unbounded); Direct prompting (one pass, no tools), plans simulated in the TravelEval sandbox of 10 Chinese cities (real rail, flight, hotel and attraction data, queuing-time model, road-network distances) over 1,150 queries; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 7 models tracked.

Top models

#ModelScore
1GPT-4o (2024-11-20)5.97
2GPT-4o Mini (2024-07-18)5.97
3Qwen 3 8B5.74
4DeepSeek V3.1 (Non-reasoning)5.72
5Gemini 2.0 Flash5.71
6GPT-5 Chat5.65
7Claude Sonnet 4.55.44

Interactive version: theaggregate.ai/benchmark?slug=traveleval-preference-weighted-experience · How It Works · Data refreshed daily, snapshot 2026-09-29.