PARROT - Result Consistency — leaderboard
Metric: Accuracy (%). Source: code4db.github.io. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | O3 Mini | 54.23 |
| 2 | O1 Preview | 48.69 |
| 3 | Claude 3.7 Sonnet | 22.74 |
| 4 | GPT-4o | 21.87 |
| 5 | Doubao-1.5-Pro | 14.29 |
Interactive version: theaggregate.ai/benchmark?slug=parrot-result-consistency · How the rankings work · Data refreshed daily, snapshot 2026-07-22.