PrinzBench: leaderboard

Private benchmark testing LLMs on legal research, analysis, and locating obscure public information. 33 questions (25 legal, 8 search) scored out of 99.

Metric: Score (x/99). Source: github.com. Status: saturation imminent. 29 models tracked.

Top models

#ModelScore
1GPT-5.6 Pro Sol91
2GPT-5.4 (xHigh)69
3GPT-5.3 Codex (High)52
4Gemini 3.1 Pro (Preview)50
5Kimi K347
6Grok 4.2043
7Gemini 3 Flash36
8Gemini 3 Pro35
9Kimi K2.5 (Thinking)35
10Muse Spark31
11GLM-5.230
12Claude Opus 4.725
13Qwen 3 Max25
14Grok 423
15DeepSeek V4 Pro23

Interactive version: theaggregate.ai/benchmark?slug=prinzbench · How It Works · Data refreshed daily, snapshot 2026-09-05.