AgroTools: leaderboard

Metric: Final answer score (0-100) in end-to-end mode (the agent plans, calls the executable tools and answers with no gold history; Lagent ReAct schema, temperature 0.1) on the AgroTools shared subset (image-grounded agricultural queries from 12 public datasets, excluding raster-processing and plot-generation queries); answers are normalized into slots by a GPT-4o extractor, closed slots scored by deterministic matching and open-text slots by a GPT-4o judge against reference key points, each query scored 0-1 and averaged (times 100, so 0-100); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 13 models tracked.

Top models

#ModelScore
1GPT-5.462.74
2GLM-4.6V52.15
3Seed 2.0 Pro48.64
4Qwen 3 VL 8B Instruct41.64
5Gemini 2.5 Pro40.68
6Qwen 3.5 9B39.19
7Qwen 3.5 35B A3B38.53
8Claude Sonnet 4.625.44
9InternVL3-8B1.61

Interactive version: theaggregate.ai/benchmark?slug=agrotools · How It Works · Data refreshed daily, snapshot 2026-10-07.