LogicSkills - Countermodel Construction: leaderboard

Metric: Accuracy (fraction times 100) on 300 language-neutral countermodel items: for an invalid argument the model gives a finite structure on a fixed small domain that makes every premise true and the conclusion false, checked by Z3; greedy decoding, responses normalized by a GPT-4o extractor and checked with the Z3 solver; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 10 models tracked.

Top models

#ModelScoreOverall rank
1O399#121
2Qwen 3 32B (Thinking)89#424 (Qwen 3 32B)
3Claude 3.7 Sonnet42#241
4Qwen2.5-Math-72B-Instruct15#708
5Phi-414#701
6Gemini 2.5 Flash13#237
7Llama 3.1 70B Instruct12#548
8GPT-4o10#333
9Llama 3.1 8B Instruct0#1018

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=logicskills-countermodel-construction · How It Works · Data refreshed daily, snapshot 2026-10-11.