AJ-Bench (LLM-as-a-Judge) - PPT: leaderboard

Metric: F1 (%) on PowerPoint GUI tasks (OSWorld; 21 tasks, 42 trajectories), averaged over three runs, with the judge reading the trajectory only (no tools or environment access), of AJ-Bench, judging whether agent trajectories succeeded against binary human-verified labels (155 tasks, 516 trajectories), default model configurations; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 13 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)76.1
2Claude Sonnet 4.575.61
3Gemini 2.5 Pro68.72
4Kimi K2 090565.53
5Grok 461.11
6GLM-4.660.82
7Claude Opus 4.559.21
8DeepSeek V3.258.38
9GPT-551.9
10Qwen 3 235B A22B45.5
11LongCat-Flash-Chat45.33
12GPT-5 Mini (Low)45.05
13GPT-5.141.9

Interactive version: theaggregate.ai/benchmark?slug=aj-bench-llm-as-a-judge-ppt · How It Works · Data refreshed daily, snapshot 2026-10-07.