Can LLMs Falsify? — leaderboard

Tests whether LLMs can generate counterexamples to falsify incorrect mathematical hypotheses - a core scientific reasoning capability.

Metric: Counterexample Rate (%). Source: falsifiers.github.io. Status: years away from saturation. 5 models tracked.

Top models

#ModelScore
1O3 Mini (High)8.9
2DeepSeek R15.8
3Claude 3.5 Sonnet4.6
4DeepSeek V32.7
5Gemini 2.0 Flash (Thinking)2.1

Interactive version: theaggregate.ai/benchmark?slug=can-llms-falsify · How the rankings work · Data refreshed daily, snapshot 2026-07-22.