Trip+ - Feasibility: leaderboard

Metric: Itinerary feasibility (0-100), a rate scaled by 100: deterministic checks of structure, grounded venues and transport, timing, opening hours, transfers and cost arithmetic, averaged over the plans that pass the response-mode gate, on multi-turn personalized travel planning with traveler profiles, request changes and environment disruptions, the same OpenAI-compatible function-calling scaffold over a fixed travel sandbox for every model, temperature 0 where supported; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 18 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)88.12
2GLM-5.184.63
3GPT-5.476.69
4Qwen 3.6 27B (Non-reasoning)75.21
5Kimi K2.674.72
6Gemma 4 31B74.27
7Qwen 3.5 27B (Non-reasoning)73.93
8Hy3-preview72.23
9Gemini 3 Flash (Preview)70.99
10Gemma 4 26B A4B67.61
11DeepSeek V4 Pro67.55
12Seed 2.0 Pro67.31
13DeepSeek V3.264.61
14GPT-5.4 Mini64.55
15MiniMax-M2.755.96

Interactive version: theaggregate.ai/benchmark?slug=trip-plus-feasibility · How It Works · Data refreshed daily, snapshot 2026-09-29.