JourneyBench (Static-Prompt Agent) - Failing Function: leaderboard

Metric: User Journey Coverage Score (0-1; a conversation scores 0 on any missing, extra or misordered tool call, otherwise its tool-parameter accuracy; GPT-4o simulated user, 40-turn limit; Failing Function scenarios; the whole SOP graph compiled into one static system prompt). Source: arxiv.org. Saturation forecast: Estimated already saturated. 4 models tracked.

Top models

#ModelScore
1GPT-4o0.51
2GPT-4o Mini0.33
3Claude 3.5 Haiku0.28
4Llama 3.3 70B Instruct0.26

Interactive version: theaggregate.ai/benchmark?slug=journeybench-static-prompt-agent-failing-function · How It Works · Data refreshed daily, snapshot 2026-09-26.