GISA - Item EM: leaderboard

Metric: Exact match (%) on the 22 item-type queries (a single answer); GISA's human-written information-seeking queries with deterministic answers (normalized TSV output compared cell by cell); LLMs run as ReAct agents with Google search (Serper) and Jina browse tools whose pages the same model summarizes, at most 30 tool calls and 8,192 output tokens per step, while commercial search systems run as offered; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 16 models tracked.

Top models

#ModelScoreOverall rank
1Kimi K2.5 (Thinking)68.18#139 (Kimi K2.5)
2Claude Sonnet 4.5 (Thinking)63.64#138 (Claude Sonnet 4.5)
3DeepSeek V3.2 (Thinking)63.64#198 (DeepSeek V3.2)
4GPT-5.2 (Thinking)63.64#105 (GPT-5.2)
5Claude Sonnet 4.559.09#138
6Qwen 3 Max (Thinking)59.09#201 (Qwen 3 Max)
7Gemini 3 Pro (High)50#77 (Gemini 3 Pro)
8GLM-4.7 (Thinking)50#185 (GLM-4.7)
9Gemini 3 Pro (Low)45.45#77 (Gemini 3 Pro)
10Qwen 3 235B A22B (Thinking)40.91#304 (Qwen 3 235B A22B)
11DeepSeek V3.2 (Non-reasoning)22.73#198 (DeepSeek V3.2)
12O4 Mini Deep Research18.18#219
13GPT-4o Search Preview13.64#445

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=gisa-item-em · How It Works · Data refreshed daily, snapshot 2026-10-11.