SciAgentBench (No Tools): leaderboard

Metric: Success rate (%, 0-100) on all 259 tasks; without tools: chain-of-thought answer only; SciAgentBench: 259 multi-step scientific tasks (1,134 sub-questions; physics 109, chemistry 81, materials 37, life sciences 32; about 65 percent with images) aggregated from existing benchmarks, kept when four frontier LLMs averaged under 50 percent and SciAgentGym could execute a verified trace; a task counts only when every sub-question is correct (strict JSON matching with 0.05 numeric tolerance, GPT-4.1 checking textual fields); temperature 0.7; higher is better. Source: arxiv.org. Saturation forecast: Around February 2028. 17 models tracked.

Top models

#ModelScoreOverall rank
1GPT-532.3#91
2Grok 4.130.4#218
3Gemini 2.5 Flash28.5#237
4O4 Mini27.8#172
5O326.6#121
6GLM-4.6V26#309
7Gemini 2.5 Pro24.8#145
8Qwen 3 VL 235B A22B (Thinking)24.4#228 (Qwen 3 VL 235B A22B)
9Qwen 3 VL 32B (Thinking)24.4#287 (Qwen 3 VL 32B)
10Qwen 3 VL 235B A22B Instruct23#264
11Qwen 3 VL 32B Instruct22.8#276
12Claude Sonnet 422.4#194
13Qwen 3 VL 8B Instruct18.4#401
14GPT-4o17.1#333
15Qwen 3 VL 4B Instruct17#506

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=sciagentbench-no-tools · How It Works · Data refreshed daily, snapshot 2026-10-11.