PrinzBench — leaderboard

Private benchmark testing LLMs on legal research, analysis, and locating obscure public information. 33 questions (25 legal, 8 search) scored out of 99.

Metric: Score (x/99). Source: github.com. Status: saturation imminent. 29 models tracked.

Top models

#ModelScore
1GPT-5.4 (xHigh)69
2GPT-5.3 Codex (High)52
3Gemini 3.1 Pro (Preview)50
4Kimi K347
5Grok 4.2043
6Gemini 3 Flash36
7Gemini 3 Pro35
8Kimi K2.5 (Thinking)35
9GLM-5.230
10Claude Opus 4.725
11Qwen 3 Max25
12Grok 4.125
13Grok 423
14DeepSeek V4 Pro23
15Kimi K2 (Thinking)22

Interactive version: theaggregate.ai/benchmark?slug=prinzbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.