LinkS2Bench: leaderboard

Metric: Average (%) of the 12 task scores (accuracy, MRA for zone counting, ACC@1s for temporal grounding) on the full LinkS2Bench (17,903 questions over 1,500 UAV videos paired with satellite images), temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around April 2028. 18 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)51.1
2Seed 1.847.8
3GPT-5.446.6
4Qwen 3.5 397B A17B45.6
5Qwen 3.5 35B A3B44.1
6Kimi K2.538.4
7GPT-5.4 Mini38.2
8Claude Sonnet 4.637.7
9Qwen 3.5 9B30.7
10Llama 4 Scout26.6
11Qwen 3.5 4B23.7
12Ministral 3 14B23.3
13Ministral 3 8B22.9
14Ministral 3 3B22.9

Interactive version: theaggregate.ai/benchmark?slug=links2bench · How It Works · Data refreshed daily, snapshot 2026-10-07.