OrchestrationBench — leaderboard

Multi-step workflow orchestration benchmark testing planning, function calling accuracy, and call rejection ability in Korean and English.

Metric: Average Score. Source: github.com. Status: saturation imminent. 17 models tracked.

Top models

#ModelScore
1Claude Opus 4.785.07
2Gemini 3 Pro (Preview)84.12
3Gemini 3 Flash (Preview)83.54
4GLM-4.783.17
5Gemini 3.1 Pro (Preview)82.8
6GPT-5.4 (2026-03-05)78.15
7GPT-5.276.05
8Qwen 3 235B A22B 2507 Instruct70.65
9Qwen 3 30B A3B 2507 Instruct66.08
10Qwen 2.5 32B Instruct66.03

Interactive version: theaggregate.ai/benchmark?slug=orchestrationbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.