T2S-Bench - Fault Localization: leaderboard

Metric: Exact match (%) on the 204 Fault Localization questions on T2S-Bench-MR (500 expert-checked multiple-choice questions over text-structure pairs from scientific papers in six domains; single- and multiple-answer items, only the text and the question given), deterministic decoding at temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 45 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro74.51#145
2Claude Sonnet 4.565.2#138
3Gemini 2.5 Flash64.22#237
4GPT-5.262.75#105
5Claude Sonnet 4 (20250514)61.27#211
6Qwen 3 32B53.43#424
7GPT-5.151.47#131
8Kimi K2 090550.98#282
9Qwen 3 235B A22B 2507 Instruct50#291
10Qwen 3 235B A22B 2507 (Thinking)50#253 (Qwen 3 235B A22B 2507)
11Claude Haiku 4.5 (20251001)49.51#251
12GPT-4.1 Mini48.04#346
13DeepSeek R1 052846.57#217
14GPT-4o44.12#333
15DeepSeek V3 (0324)43.63#332

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=t2s-bench-fault-localization · How It Works · Data refreshed daily, snapshot 2026-10-11.