GISA - List Order Score: leaderboard

Metric: Order score (%) on the 48 list-type queries: twice the matching elements over the total length of both lists (difflib SequenceMatcher ratio, times 100); GISA's human-written information-seeking queries with deterministic answers (normalized TSV output compared cell by cell); LLMs run as ReAct agents with Google search (Serper) and Jina browse tools whose pages the same model summarizes, at most 30 tool calls and 8,192 output tokens per step, while commercial search systems run as offered; higher is better. Source: arxiv.org. Saturation forecast: Around May 2027. 16 models tracked.

Top models

#ModelScoreOverall rank
1Qwen 3 Max (Thinking)64.08#201 (Qwen 3 Max)
2DeepSeek V3.2 (Thinking)60.41#198 (DeepSeek V3.2)
3Gemini 3 Pro (High)60.12#77 (Gemini 3 Pro)
4Claude Sonnet 4.557.78#138
5Claude Sonnet 4.5 (Thinking)56.42#138 (Claude Sonnet 4.5)
6Gemini 3 Pro (Low)56.37#77 (Gemini 3 Pro)
7DeepSeek V3.2 (Non-reasoning)55.45#198 (DeepSeek V3.2)
8GPT-5.2 (Thinking)53.17#105 (GPT-5.2)
9O4 Mini Deep Research52.59#219
10GLM-4.7 (Thinking)50.97#185 (GLM-4.7)
11Kimi K2.5 (Thinking)48.81#139 (Kimi K2.5)
12GPT-4o Search Preview36#445
13Qwen 3 235B A22B (Thinking)35.96#304 (Qwen 3 235B A22B)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=gisa-list-order-score · How It Works · Data refreshed daily, snapshot 2026-10-11.