ClinSQL - Laboratory Results Analysis: leaderboard

Metric: Execution Score (%). Source: arxiv.org. Saturation forecast: Around December 2026. 22 models tracked.

Top models

#ModelScore
1DeepSeek R1 052866.91
2GPT-5 Mini65.09
3Gemini 2.5 Pro63.66
4GPT-5 Chat57.89
5Qwen 3 235B A22B 2507 Instruct55.94
6Gemini 2.5 Flash54.4
7DeepSeek V3.153.66
8Qwen 3 Coder 480B A35B Instruct49.78
9GPT-5 Nano49.02
10O4 Mini (2025-04-16)48.45
11GPT-4.148.39
12Grok 4 Fast (Reasoning)45.5
13Qwen 3 Next 80B A3B Instruct42.33
14Llama 4 Maverick Instruct37.66
15Mistral Medium 332.22

Interactive version: theaggregate.ai/benchmark?slug=clinsql-laboratory-results-analysis · How It Works · Data refreshed daily, snapshot 2026-09-25.