T2S-Bench - Link Extraction: leaderboard

Metric: Link F1 (%) on T2S-Bench-E2E (87 human-verified text-structure pairs): given the text and every node, the model outputs the links as JSON, scored by F1 against the reference links, temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 45 models tracked.

Top models

#ModelScoreOverall rank
1Claude Sonnet 4.586.91#138
2Gemini 2.5 Pro84.32#145
3Claude Sonnet 4 (20250514)84.07#211
4DeepSeek R1 052880.31#217
5GPT-5.179.44#131
6Claude Haiku 4.5 (20251001)79.33#251
7DeepSeek V3.278.69#198
8DeepSeek V3 (0324)78.17#332
9DeepSeek V3.177.77#260
10GPT-5.277.76#105
11Qwen 3 Next 80B A3B (Thinking)77.64#306 (Qwen 3 Next 80B A3B)
12Qwen 3 Next 80B A3B Instruct76.11#352
13Qwen 3 235B A22B 2507 (Thinking)76.11#253 (Qwen 3 235B A22B 2507)
14Claude 3 Haiku (20240307)75.51#759
15Gemini 2.5 Flash75.1#237

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=t2s-bench-link-extraction · How It Works · Data refreshed daily, snapshot 2026-10-11.