JourneyBench (Static-Prompt Agent) - Correct Context: leaderboard

Metric: User Journey Coverage Score (0-1; a conversation scores 0 on any missing, extra or misordered tool call, otherwise its tool-parameter accuracy; GPT-4o simulated user, 40-turn limit; Correct Context scenarios; the whole SOP graph compiled into one static system prompt). Source: arxiv.org. Saturation forecast: Estimated already saturated. 4 models tracked.

Top models

#ModelScore
1GPT-4o0.87
2GPT-4o Mini0.72
3Llama 3.3 70B Instruct0.24
4Claude 3.5 Haiku0.23

Interactive version: theaggregate.ai/benchmark?slug=journeybench-static-prompt-agent-correct-context · How It Works · Data refreshed daily, snapshot 2026-09-26.