owl-alpha: benchmark results
Provider: Other. Access: API.
Unified ELO 1707 ± 81, rank #317 of 2656 rated models, from 10 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BLXBench | 83.6 | Score (self-reported) | 87.5 |
| SvelteBench | 92.2 | Average pass@1 (%) | 68.9 |
| THOR Finding Triage - CW% | 56 | Confidence-weighted classification score (%) | 46.2 |
| THOR Finding Triage - False Review Load | 47.4 | False positives not suppressed (%, lower is better) | 35 |
| THOR Finding Triage - Balanced OTS | 53 | Class-balanced operational triage score (%) | 30 |
| THOR Finding Triage - Critical Miss Rate | 3.6 | True positives suppressed as false positives (%, lower is be | 28.7 |
| Tinybird AI SQL Benchmark - Exactness | 41.26 | Result exactness vs human reference queries (0-100) | 23.9 |
| SpeechMap Compliance | 35.1 | % Requests Completed | 13 |
| Tinybird AI SQL Benchmark - First-Attempt Success Rate | 62 | Questions answered with a valid query on the first attempt ( | 10 |
| Tinybird AI SQL Benchmark - Success Rate | 66 | Questions answered with a valid query within 3 retries (%) | 6.9 |
Interactive version: theaggregate.ai/model?slug=owl-alpha · How It Works · Data refreshed daily, snapshot 2026-09-19.