JourneyBench (Dynamic-Prompt Agent) - Loan Application: leaderboard

Metric: User Journey Coverage Score (0-1; a conversation scores 0 on any missing, extra or misordered tool call, otherwise its tool-parameter accuracy; GPT-4o simulated user, 40-turn limit; Loan Application SOP graph; the authors' orchestrator walks the SOP graph node by node and swaps the prompt and tools at each transition). Source: arxiv.org. Saturation forecast: Estimated already saturated. 4 models tracked.

Top models

#ModelScore
1GPT-4o0.78
2GPT-4o Mini0.62
3Claude 3.5 Haiku0.62
4Llama 3.3 70B Instruct0.33

Interactive version: theaggregate.ai/benchmark?slug=journeybench-dynamic-prompt-agent-loan-application · How It Works · Data refreshed daily, snapshot 2026-09-26.