AgentJudgeBench (Hard, With Ground Truth) - SmolLM3-3B Outputs: leaderboard

Metric: Judge alignment (%; agreement of the judge's per-metric verdicts (tool selection, parameter structure, call sequence, query coverage) with the programmatic reference on the hard (most ambiguous) rewrite of 3,808 BFCL-style tool-calling records over six workflow-DAG topologies; judged outputs from SmolLM3-3B at temperature 0; the judge sees the ground-truth tool calls). Source: arxiv.org. Saturation forecast: Around January 2028. 6 models tracked.

Top models

#ModelScore
1GPT-OSS-120B77.8
2GPT-OSS-20B77.4
3QwQ-32B77.2
4GPT-5.476.9
5Claude Sonnet 4.574.2
6Gemini 2.5 Pro (Preview 05-06)73.4

Interactive version: theaggregate.ai/benchmark?slug=agentjudgebench-hard-with-ground-truth-smollm3-3b-outputs · How It Works · Data refreshed daily, snapshot 2026-09-29.