JourneyBench (Static-Prompt Agent) - Telecommunications: leaderboard

Metric: User Journey Coverage Score (0-1; a conversation scores 0 on any missing, extra or misordered tool call, otherwise its tool-parameter accuracy; GPT-4o simulated user, 40-turn limit; Telecommunications SOP graph; the whole SOP graph compiled into one static system prompt). Source: arxiv.org. Saturation forecast: Estimated already saturated. 4 models tracked.

Top models

#ModelScore
1GPT-4o0.42
2GPT-4o Mini0.3
3Llama 3.3 70B Instruct0.12
4Claude 3.5 Haiku0.12

Interactive version: theaggregate.ai/benchmark?slug=journeybench-static-prompt-agent-telecommunications · How It Works · Data refreshed daily, snapshot 2026-09-26.