AgentJudgeBench (Hard, With Ground Truth) - Llama-3.1-8B-Instruct Outputs: leaderboard
Metric: Judge alignment (%; agreement of the judge's per-metric verdicts (tool selection, parameter structure, call sequence, query coverage) with the programmatic reference on the hard (most ambiguous) rewrite of 3,808 BFCL-style tool-calling records over six workflow-DAG topologies; judged outputs from Llama-3.1-8B-Instruct at temperature 0; the judge sees the ground-truth tool calls). Source: arxiv.org. Saturation forecast: Around December 2026. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | QwQ-32B | 84.1 |
| 2 | GPT-OSS-120B | 83.6 |
| 3 | GPT-OSS-20B | 81.2 |
| 4 | GPT-5.4 | 80.6 |
| 5 | Claude Sonnet 4.5 | 80.3 |
| 6 | Gemini 2.5 Pro (Preview 05-06) | 78.7 |
Interactive version: theaggregate.ai/benchmark?slug=agentjudgebench-hard-with-ground-truth-llama-3-1-8b-instruct-outputs · How It Works · Data refreshed daily, snapshot 2026-09-29.