PluginEval - Over-Recall: leaderboard

Metric: Over-recall rate (%; share of the 500 adversarial negative queries on which the model wrongly calls a tool; lower is better; 3,000 human-verified Chinese function-calling queries over 54 plugins and 63 tools with real API execution (2,500 positive queries requiring grounded calls, 500 adversarial negatives requiring abstention; easy/medium/hard 0.4/0.4/0.2); GPT-5.2 judge anchored to gold annotations). Source: arxiv.org. Saturation forecast: Around June 2027. 5 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)18.55
2Qwen 3 235B A22B24.22
3DeepSeek V4 Pro28.91
4Claude Opus 4.629.1
5GPT-5.433.2

Interactive version: theaggregate.ai/benchmark?slug=plugineval-over-recall · How It Works · Data refreshed daily, snapshot 2026-09-29.