LibEvoBench - Stable APIs: leaderboard

Metric: Average task performance (0-100) on stable APIs (identical signature in every covered version): mean of four task metrics (API calling with version constraint, exact match; API calling parameter recall given name and version; API identification from a redacted docstring, exact match; signature recall F1) averaged over PyTorch, NumPy and SciPy versions, temperature 0 where supported and no thinking budget; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.

Top models

#ModelScore
1GPT-5.592.3
2GPT-5.492
3Claude Sonnet 4.689.3
4Claude Sonnet 487.8
5Gemini 3 Flash (Minimal)87.8
6GPT-4.184.3
7GPT-5.183.8
8GPT-583
9Gemini 2.5 Flash (Non-reasoning)81.6
10Qwen 3.5 397B A17B (Non-reasoning)80.2
11Gemini 2.0 Flash79.7
12Qwen 3.5 122B A10B (Non-reasoning)70.5
13Qwen 3.5 35B A3B (Non-reasoning)63.3

Interactive version: theaggregate.ai/benchmark?slug=libevobench-stable-apis · How It Works · Data refreshed daily, snapshot 2026-09-29.