T2S-Bench - Counterfactual Reasoning: leaderboard

Metric: Exact match (%) on the 107 Counterfactual Reasoning questions on T2S-Bench-MR (500 expert-checked multiple-choice questions over text-structure pairs from scientific papers in six domains; single- and multiple-answer items, only the text and the question given), deterministic decoding at temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 45 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro90.65#145
2Qwen 3 32B89.72#424
3Claude Sonnet 4 (20250514)89.72#211
4Claude Sonnet 4.588.79#138
5Claude Haiku 4.5 (20251001)85.98#251
6GPT-5.284.11#105
7Kimi K2 090584.11#282
8GPT-5.183.18#131
9Mistral Small 3.280.37#587
10GPT-4o79.44#333
11Gemini 2.0 Flash Lite79.44#438
12Gemini 2.5 Flash78.5#237
13DeepSeek V3.278.5#198
14Qwen 3 235B A22B 2507 Instruct78.5#291
15DeepSeek V3.178.5#260

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=t2s-bench-counterfactual-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-11.