Trip+ - Response Mode Accuracy: leaderboard

Metric: Response-mode accuracy (0-100), a rate scaled by 100: whether the agent answers a turn with a plan, a clarification or a no-solution reply as the hidden turn state expects, averaged over all turns, on multi-turn personalized travel planning with traveler profiles, request changes and environment disruptions, the same OpenAI-compatible function-calling scaffold over a fixed travel sandbox for every model, temperature 0 where supported; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 18 models tracked.

Top models

#ModelScore
1Gemma 4 31B90.35
2Hy3-preview89.3
3GLM-5.187.02
4Gemini 3 Flash (Preview)86.67
5Gemini 3.1 Pro (Preview)85.44
6Qwen 3.5 27B (Non-reasoning)84.91
7Kimi K2.684.39
8Gemma 4 26B A4B83.86
9Seed 2.0 Pro82.46
10DeepSeek V4 Pro79.12
11GPT-5.4 Mini78.07
12MiniMax-M2.777.72
13DeepSeek V3.277.02
14GPT-5.475.44
15Qwen 3.6 35B A3B (Non-reasoning)69.12

Interactive version: theaggregate.ai/benchmark?slug=trip-plus-response-mode-accuracy · How It Works · Data refreshed daily, snapshot 2026-09-29.