Trip+ - Hard Constraints: leaderboard

Metric: Hard-constraint satisfaction (0-100), a rate scaled by 100: explicit user requirements (dates, destinations, party size, budget, required lodging, dining and transport) met by rule checks, averaged over the plans that pass the response-mode gate, on multi-turn personalized travel planning with traveler profiles, request changes and environment disruptions, the same OpenAI-compatible function-calling scaffold over a fixed travel sandbox for every model, temperature 0 where supported; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 18 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)90.47
2DeepSeek V4 Pro83.75
3GPT-5.481.1
4Seed 2.0 Pro80.06
5Kimi K2.679.84
6GLM-5.179.73
7Gemma 4 31B76.32
8Qwen 3.5 27B (Non-reasoning)72.56
9Qwen 3.6 27B (Non-reasoning)72.25
10Gemini 3 Flash (Preview)71.79
11DeepSeek V3.271.62
12Gemma 4 26B A4B69.25
13Hy3-preview64.76
14MiniMax-M2.763.26
15GPT-5.4 Mini61.26

Interactive version: theaggregate.ai/benchmark?slug=trip-plus-hard-constraints · How It Works · Data refreshed daily, snapshot 2026-09-29.