GPT-6.1 Pro Sol: benchmark results
Provider: OpenAI. Access: API.
Unified ELO 1912 ± 29, rank #12 of 1636 rated models, from 19 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| MineBench | 2212 | Elo Rating | 98.6 |
| Arabic Broad Leaderboard - MMLU | 9.92 | Average Score (0-10) | 98.3 |
| Arabic Broad Leaderboard - Translation (incl Dialects) | 8.06 | Average Score (0-10) | 96.2 |
| AIMultiple - FinanceReasoning (Hard) | 90.34 | Accuracy (%, 238 hard FinanceReasoning questions) | 95.5 |
| Arabic Broad Leaderboard | 9.05 | Average Score (0-10) | 91.6 |
| Arabic Broad Leaderboard - Trust & Safety | 9.67 | Average Score (0-10) | 91.6 |
| PRISM (1C:Enterprise) - Platform Tasks (B) | 90 | 1C platform tasks fully solved, run in headless 1C (%) | 91.5 |
| AIMultiple - Text-to-SQL (Adjudicated) | 79.7 | Execution match with jury-credited equivalents and broken go | 87.7 |
| Arabic Broad Leaderboard - RAG QA | 8.39 | Average Score (0-10) | 84 |
| AIMultiple - Text-to-SQL (Strict Execution Match) | 51.3 | Execution match against BIRD gold (%, over correctly routed | 83.6 |
| PRISM (1C:Enterprise) - Algorithmic Tasks (A) | 93 | Algorithmic BSL tasks fully solved, run in OneScript (%) | 80.5 |
| Tinybird AI SQL Benchmark - Exactness | 53.54 | Result exactness vs human reference queries (0-100) | 79 |
Interactive version: theaggregate.ai/model?slug=gpt-6-1-pro-sol · How It Works · Data refreshed daily, snapshot 2026-10-11.