Gorilla Benchmark API Bench: leaderboard

1,645 ML-model API calls (925 HuggingFace, 626 TensorFlow Hub, 94 TorchHub) paired with 10 synthetic instructions each, from UC Berkeley (2023); generated calls scored by AST sub-tree matching.

Source: arxiv.org.

Interactive version: theaggregate.ai/benchmark?slug=gorilla-benchmark-api-bench · How It Works · Data refreshed daily, snapshot 2026-09-05.