T2S-Bench - Node Extraction: leaderboard

Metric: Node score (%) on T2S-Bench-E2E (87 human-verified text-structure pairs): given the text and every link, the model labels the nodes, scored by mean semantic similarity to the reference node labels, temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 45 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro58.09#145
2Claude Sonnet 4.555.97#138
3Claude Sonnet 4 (20250514)54.11#211
4GPT-5.250.57#105
5Qwen 3 235B A22B 2507 Instruct49.39#291
6DeepSeek R1 052849.24#217
7DeepSeek V3 (0324)47.32#332
8Claude Haiku 4.5 (20251001)47.06#251
9DeepSeek V3.246.98#198
10Gemini 2.5 Flash46.9#237
11Qwen 3 14B46.85#524
12DeepSeek V3.146.59#260
13Qwen 3 235B A22B 2507 (Thinking)45.97#253 (Qwen 3 235B A22B 2507)
14Mistral Small 3.245.74#587
15Mistral Large 2 (Nov) Instruct (2411)45.71#458

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=t2s-bench-node-extraction · How It Works · Data refreshed daily, snapshot 2026-10-11.