PatRe - Office Action Quality (OA-RO): leaderboard

Metric: Overall office action quality on a 1-10 scale: mean of the soundness, clarity, constructiveness, completeness and legal-style ratings of a Gemini-3.1-Flash-Lite judge at temperature 0, office action generation under reference oracle: the model also receives the references the examiner and applicant cited and must pick the relevant ones, over PatRe (480 recent USPTO patent examination histories across all eight IPC sections, with office actions, rebuttals, claim versions and cited references), temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around June 2028. 10 models tracked.

Top models

#ModelScore
1GPT-5 Mini4.89
2Qwen 3.5 27B4.37
3Gemini 2.5 Flash4.36
4DeepSeek V3.24.34
5Qwen 3.5 9B4.11
6Gemma 3 27B (IT)3.65
7GPT-4o Mini3.61
8Gemma 3 12B (IT)3.59
9Llama 3.3 70B Instruct3.4
10Llama 3.1 8B Instruct3.07

Interactive version: theaggregate.ai/benchmark?slug=patre-office-action-quality-oa-ro · How It Works · Data refreshed daily, snapshot 2026-10-07.