LMGame-Bench Ace Attorney — leaderboard

LLM game-playing benchmark: Ace Attorney testing logical deduction and evidence-based reasoning in courtroom scenarios.

Metric: Score. Source: huggingface.co. Status: saturation imminent. 17 models tracked.

Top models

#ModelScore
1O3 (2025-04-16)16
2O1 (2024-12-17)16
3GPT-5 (High)9
4Claude 3.7 Sonnet (20250219)7
5Gemini 2.5 Pro (Preview 05-06)7
6Claude Opus 4 (20250514)6
7O4 Mini (2025-04-16)4
8Gemini 2.5 Flash (Preview 04-17)4
9Claude Sonnet 4 (20250514)3.7
10Claude 3.5 Sonnet (20241022)2
11GPT-4.12
12DeepSeek R10
13GPT-4o (2024-11-20)0
14Llama 4 Maverick Instruct FP80
15O1 Mini (2024-09-12)0

Interactive version: theaggregate.ai/benchmark?slug=lmgame-bench-ace-attorney · How the rankings work · Data refreshed daily, snapshot 2026-07-22.