GISA - Table EM: leaderboard
Metric: Exact match (%) on the 253 table-type queries (a table with a predefined schema, every cell right); GISA's human-written information-seeking queries with deterministic answers (normalized TSV output compared cell by cell); LLMs run as ReAct agents with Google search (Serper) and Jina browse tools whose pages the same model summarizes, at most 30 tool calls and 8,192 output tokens per step, while commercial search systems run as offered; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 16 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Claude Sonnet 4.5 (Thinking) | 13.04 | #138 (Claude Sonnet 4.5) |
| 2 | Qwen 3 Max (Thinking) | 10.67 | #201 (Qwen 3 Max) |
| 3 | Claude Sonnet 4.5 | 9.49 | #138 |
| 4 | GPT-5.2 (Thinking) | 9.49 | #105 (GPT-5.2) |
| 5 | Gemini 3 Pro (High) | 8.7 | #77 (Gemini 3 Pro) |
| 6 | GLM-4.7 (Thinking) | 8.3 | #185 (GLM-4.7) |
| 7 | Kimi K2.5 (Thinking) | 7.91 | #139 (Kimi K2.5) |
| 8 | Gemini 3 Pro (Low) | 7.11 | #77 (Gemini 3 Pro) |
| 9 | DeepSeek V3.2 (Non-reasoning) | 6.72 | #198 (DeepSeek V3.2) |
| 10 | DeepSeek V3.2 (Thinking) | 6.32 | #198 (DeepSeek V3.2) |
| 11 | GPT-4o Search Preview | 4.74 | #445 |
| 12 | Qwen 3 235B A22B (Thinking) | 4.35 | #304 (Qwen 3 235B A22B) |
| 13 | O4 Mini Deep Research | 3.56 | #219 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=gisa-table-em · How It Works · Data refreshed daily, snapshot 2026-10-11.