Gorilla Benchmark API Bench: leaderboard
1,645 ML-model API calls (925 HuggingFace, 626 TensorFlow Hub, 94 TorchHub) paired with 10 synthetic instructions each, from UC Berkeley (2023); generated calls scored by AST sub-tree matching.
Source: arxiv.org.
Interactive version: theaggregate.ai/benchmark?slug=gorilla-benchmark-api-bench · How It Works · Data refreshed daily, snapshot 2026-09-05.