LogicSkills - Validity Assessment (Carroll): leaderboard

Metric: Accuracy (fraction times 100) on 300 validity-assessment items in a Carroll-style nonce-word language: given three to five premises and six candidate conclusions from the two-variable fragment of first-order logic, the model returns the set of conclusions that follow (set equality with the Z3-verified gold set); greedy decoding, responses normalized by a GPT-4o extractor and checked with the Z3 solver; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 10 models tracked.

Top models

#ModelScoreOverall rank
1O399#121
2Gemini 2.5 Flash96#237
3Claude 3.7 Sonnet96#241
4Qwen 3 32B (Thinking)95#424 (Qwen 3 32B)
5GPT-4o87#333
6Llama 3.1 70B Instruct84#548
7Phi-480#701
8Qwen2.5-Math-72B-Instruct80#708
9Llama 3.1 8B Instruct46#1018

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=logicskills-validity-assessment-carroll · How It Works · Data refreshed daily, snapshot 2026-10-11.