PatRe - Rebuttal Quality: leaderboard

Metric: Overall rebuttal quality on a 1-10 scale: mean of the soundness, clarity, constructiveness, completeness and legal-style ratings of a Gemini-3.1-Flash-Lite judge at temperature 0, for applicant rebuttals to the office actions of PatRe (480 recent USPTO patent examination histories across all eight IPC sections, with office actions, rebuttals, claim versions and cited references), temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 10 models tracked.

Top models

#ModelScore
1GPT-5 Mini9.18
2DeepSeek V3.28.37
3Gemini 2.5 Flash8.34
4Qwen 3.5 27B8.29
5Qwen 3.5 9B7.09
6Gemma 3 27B (IT)5.47
7Llama 3.3 70B Instruct4.59
8Gemma 3 12B (IT)4.58
9GPT-4o Mini4.5
10Llama 3.1 8B Instruct3.71

Interactive version: theaggregate.ai/benchmark?slug=patre-rebuttal-quality · How It Works · Data refreshed daily, snapshot 2026-10-07.