LogicSkills - Formal Symbolization (Carroll): leaderboard

Metric: Accuracy (fraction times 100) on 300 formal-symbolization items in a Carroll-style nonce-word language: the model translates one sentence into a two-variable first-order formula with a fixed symbol key, correct when Z3 finds it logically equivalent to the target; greedy decoding, responses normalized by a GPT-4o extractor and checked with the Z3 solver; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 10 models tracked.

Top models

#ModelScoreOverall rank
1O397#121
2Qwen 3 32B (Thinking)82#424 (Qwen 3 32B)
3Qwen2.5-Math-72B-Instruct80#708
4Gemini 2.5 Flash76#237
5GPT-4o74#333
6Claude 3.7 Sonnet71#241
7Llama 3.1 70B Instruct66#548
8Phi-451#701
9Llama 3.1 8B Instruct14#1018

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=logicskills-formal-symbolization-carroll · How It Works · Data refreshed daily, snapshot 2026-10-11.