GISA - Table Row F1: leaderboard

Metric: Row-level F1 (%) between predicted and gold table rows on the 253 table-type queries; GISA's human-written information-seeking queries with deterministic answers (normalized TSV output compared cell by cell); LLMs run as ReAct agents with Google search (Serper) and Jina browse tools whose pages the same model summarizes, at most 30 tool calls and 8,192 output tokens per step, while commercial search systems run as offered; higher is better. Source: arxiv.org. Saturation forecast: Around June 2028. 16 models tracked.

Top models

#ModelScoreOverall rank
1Claude Sonnet 4.5 (Thinking)49.92#138 (Claude Sonnet 4.5)
2Qwen 3 Max (Thinking)48.48#201 (Qwen 3 Max)
3Claude Sonnet 4.547.85#138
4Gemini 3 Pro (High)47.01#77 (Gemini 3 Pro)
5Gemini 3 Pro (Low)45.93#77 (Gemini 3 Pro)
6Kimi K2.5 (Thinking)45.19#139 (Kimi K2.5)
7DeepSeek V3.2 (Non-reasoning)44.14#198 (DeepSeek V3.2)
8GLM-4.7 (Thinking)43.97#185 (GLM-4.7)
9DeepSeek V3.2 (Thinking)43.44#198 (DeepSeek V3.2)
10GPT-5.2 (Thinking)43.04#105 (GPT-5.2)
11O4 Mini Deep Research36.78#219
12GPT-4o Search Preview29.59#445
13Qwen 3 235B A22B (Thinking)28.32#304 (Qwen 3 235B A22B)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=gisa-table-row-f1 · How It Works · Data refreshed daily, snapshot 2026-10-11.