T2S-Bench: leaderboard

Metric: Exact match (%) over all four question categories on T2S-Bench-MR (500 expert-checked multiple-choice questions over text-structure pairs from scientific papers in six domains; single- and multiple-answer items, only the text and the question given), deterministic decoding at temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 45 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro81.4#145
2Claude Sonnet 4.576.8#138
3Claude Sonnet 4 (20250514)75.6#211
4Gemini 2.5 Flash72.2#237
5GPT-5.271.8#105
6Qwen 3 32B69.4#424
7GPT-5.167.8#131
8Claude Haiku 4.5 (20251001)67.4#251
9Kimi K2 090567#282
10Qwen 3 235B A22B 2507 Instruct62.8#291
11GPT-4o61.8#333
12Gemini 2.0 Flash61.2#331
13Qwen 3 235B A22B 2507 (Thinking)60.8#253 (Qwen 3 235B A22B 2507)
14DeepSeek V3 (0324)60.6#332
15DeepSeek V3.160.4#260

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=t2s-bench · How It Works · Data refreshed daily, snapshot 2026-10-11.