T2S-Bench - Functional Mapping: leaderboard

Metric: Exact match (%) on the 69 Functional Mapping questions on T2S-Bench-MR (500 expert-checked multiple-choice questions over text-structure pairs from scientific papers in six domains; single- and multiple-answer items, only the text and the question given), deterministic decoding at temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 45 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro89.86#145
2Claude Sonnet 4 (20250514)81.16#211
3Claude Sonnet 4.579.71#138
4GPT-5.178.26#131
5Qwen 3 235B A22B 2507 Instruct78.26#291
6Gemini 2.5 Flash76.81#237
7Gemini 2.0 Flash76.81#331
8DeepSeek V3.173.91#260
9DeepSeek V3 (0324)73.91#332
10DeepSeek V3.272.46#198
11Claude Haiku 4.5 (20251001)72.46#251
12Qwen 3 235B A22B 2507 (Thinking)72.46#253 (Qwen 3 235B A22B 2507)
13DeepSeek R1 052872.46#217
14Kimi K2 090572.46#282
15GPT-5.269.57#105

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=t2s-bench-functional-mapping · How It Works · Data refreshed daily, snapshot 2026-10-11.