AgroTools: leaderboard
Metric: Final answer score (0-100) in end-to-end mode (the agent plans, calls the executable tools and answers with no gold history; Lagent ReAct schema, temperature 0.1) on the AgroTools shared subset (image-grounded agricultural queries from 12 public datasets, excluding raster-processing and plot-generation queries); answers are normalized into slots by a GPT-4o extractor, closed slots scored by deterministic matching and open-text slots by a GPT-4o judge against reference key points, each query scored 0-1 and averaged (times 100, so 0-100); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 | 62.74 |
| 2 | GLM-4.6V | 52.15 |
| 3 | Seed 2.0 Pro | 48.64 |
| 4 | Qwen 3 VL 8B Instruct | 41.64 |
| 5 | Gemini 2.5 Pro | 40.68 |
| 6 | Qwen 3.5 9B | 39.19 |
| 7 | Qwen 3.5 35B A3B | 38.53 |
| 8 | Claude Sonnet 4.6 | 25.44 |
| 9 | InternVL3-8B | 1.61 |
Interactive version: theaggregate.ai/benchmark?slug=agrotools · How It Works · Data refreshed daily, snapshot 2026-10-07.