LogicSkills - Validity Assessment (English): leaderboard

Metric: Accuracy (fraction times 100) on 300 validity-assessment items in controlled English: given three to five premises and six candidate conclusions from the two-variable fragment of first-order logic, the model returns the set of conclusions that follow (set equality with the Z3-verified gold set); greedy decoding, responses normalized by a GPT-4o extractor and checked with the Z3 solver; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 10 models tracked.

Top models

#ModelScoreOverall rank
1O399#121
2Qwen 3 32B (Thinking)97#424 (Qwen 3 32B)
3Gemini 2.5 Flash94#237
4Claude 3.7 Sonnet94#241
5Llama 3.1 70B Instruct86#548
6GPT-4o85#333
7Qwen2.5-Math-72B-Instruct80#708
8Phi-473#701
9Llama 3.1 8B Instruct45#1018

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=logicskills-validity-assessment-english · How It Works · Data refreshed daily, snapshot 2026-10-11.