TerraLogic: leaderboard

Metric: Final answer accuracy (%, AnsAcc) over the 545 TerraLogic optical, SAR and infrared remote-sensing tasks, judged by GPT-4o-mini with a rubric against the ground-truth answer and its key arguments; each model drives the HieraPlan tool agent under a ReAct-style protocol in OpenCompass; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 12 models tracked.

Top models

#ModelScore
1GPT-4o25.7
2GPT-3.524.4
3Llama 3 70B Instruct22.75
4Qwen 2.5 7B Instruct22.57
5DeepSeek V321.28
6Qwen 2.5 32B Instruct19.08
7InternLM3-8B-Instruct18.9
8Mistral 7B Instruct (v0.2)13.21
9Gemini 2.5 Flash12.04
10Yi-1.5-6B-Chat10.83
11Phi-3 Mini 4K Instruct3.67

Interactive version: theaggregate.ai/benchmark?slug=terralogic · How It Works · Data refreshed daily, snapshot 2026-09-29.