Trip+ - Request Fulfillment: leaderboard

Metric: Request fulfillment (0-100), a rate scaled by 100: whether the response incorporates the new or revised constraints, preferences and environment conditions of the turn, averaged over all turns, on multi-turn personalized travel planning with traveler profiles, request changes and environment disruptions, the same OpenAI-compatible function-calling scaffold over a fixed travel sandbox for every model, temperature 0 where supported; higher is better. Source: arxiv.org. Saturation forecast: Around February 2027. 18 models tracked.

Top models

#ModelScore
1Gemma 4 31B74.65
2Gemini 3.1 Pro (Preview)73.82
3GLM-5.172.14
4Gemini 3 Flash (Preview)68.85
5Seed 2.0 Pro68.11
6Kimi K2.667.27
7Hy3-preview67.08
8DeepSeek V4 Pro66.57
9Qwen 3.5 27B (Non-reasoning)63.09
10Gemma 4 26B A4B59.89
11MiniMax-M2.757.66
12GPT-5.455.85
13DeepSeek V3.254.18
14GPT-5.4 Mini48.79
15GLM-4 32B46.56

Interactive version: theaggregate.ai/benchmark?slug=trip-plus-request-fulfillment · How It Works · Data refreshed daily, snapshot 2026-09-29.