GPT-6.1 Sol (Low): benchmark results
Provider: OpenAI. Released 2026-09-29. Access: API.
Unified ELO 1805 ± 17, rank #22 of 2066 rated models, from 74 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EnigmaForge | 84.5 | Task success (%) over 600 procedurally generated story insta | 100 |
| EnigmaForge - Fact F1 | 98.5 | World-fact recovery F1 (%) over 600 instances, all items | 100 |
| EnigmaForge - Intuition | 83.8 | Task success (%) on implicit-condition instances, where the | 100 |
| AA-Omniscience Index - Software Engineering (SWE) - HTML | 92 | Omniscience Index | 99.7 |
| AA-Omniscience Index - Software Engineering (SWE) - Rust | 86 | Omniscience Index | 98.9 |
| AA-Omniscience Index - Software Engineering (SWE) - TypeScript | 93.33 | Omniscience Index | 98.8 |
| AA Long Context Reasoning | 84 | Accuracy (%) | 98 |
| AA-Omniscience Index - Software Engineering (SWE) - Julia | 84 | Omniscience Index | 98 |
| AA-Omniscience Index - Software Engineering (SWE) - Java | 68 | Omniscience Index | 97.9 |
| AA-Omniscience Index - Business | 29.9 | Omniscience Index | 97.6 |
| AA-Omniscience Index - Software Engineering (SWE) - Python | 84.5 | Omniscience Index | 97.2 |
| AA-Omniscience Index - Software Engineering (SWE) - JavaScript | 81.82 | Omniscience Index | 97.1 |
Interactive version: theaggregate.ai/model?slug=gpt-6-1-sol-low · How It Works · Data refreshed daily, snapshot 2026-10-01.