JourneyBench (Static-Prompt Agent): leaderboard

Metric: User Journey Coverage Score (0-1; a conversation scores 0 on any missing, extra or misordered tool call, otherwise its tool-parameter accuracy; GPT-4o simulated user, 40-turn limit; mean of the three domains; the whole SOP graph compiled into one static system prompt). Source: arxiv.org. Saturation forecast: Estimated already saturated. 4 models tracked.

Top models

#ModelScore
1GPT-4o0.56
2GPT-4o Mini0.44
3Claude 3.5 Haiku0.25
4Llama 3.3 70B Instruct0.25

Interactive version: theaggregate.ai/benchmark?slug=journeybench-static-prompt-agent · How It Works · Data refreshed daily, snapshot 2026-09-26.