PLSQLBench: leaderboard

Metric: Mean test pass@1 (%; mean of the three held-out test sets: MBPP+-PLSQL (308 tasks), Spider2 single-turn test (103) and Spider2 multi-turn test (63 conversations); direct generation of Oracle PL/SQL program units at temperature 0, executed on Oracle Autonomous Database 23ai; per-task fraction of unit tests passed by one solution, averaged over tasks). Source: arxiv.org. Saturation forecast: Around 2028. 8 models tracked.

Top models

#ModelScore
1GPT-5.464.96
2Claude Opus 4.863.49
3Gemma 4 31B62.26
4GPT-5.6 Sol61.71
5GPT-5.4 Mini58.6
6Grok 4.356.92
7Gemini 2.5 Flash Lite50.39
8Llama 4 Maverick46.8

Interactive version: theaggregate.ai/benchmark?slug=plsqlbench · How It Works · Data refreshed daily, snapshot 2026-09-26.