owl-alpha: benchmark results

Provider: Other. Access: API.

Unified ELO 1707 ± 81, rank #317 of 2656 rated models, from 10 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BLXBench83.6Score (self-reported)87.5
SvelteBench92.2Average pass@1 (%)68.9
THOR Finding Triage - CW%56Confidence-weighted classification score (%)46.2
THOR Finding Triage - False Review Load47.4False positives not suppressed (%, lower is better)35
THOR Finding Triage - Balanced OTS53Class-balanced operational triage score (%)30
THOR Finding Triage - Critical Miss Rate3.6True positives suppressed as false positives (%, lower is be28.7
Tinybird AI SQL Benchmark - Exactness41.26Result exactness vs human reference queries (0-100)23.9
SpeechMap Compliance35.1% Requests Completed13
Tinybird AI SQL Benchmark - First-Attempt Success Rate62Questions answered with a valid query on the first attempt (10
Tinybird AI SQL Benchmark - Success Rate66Questions answered with a valid query within 3 retries (%)6.9

Interactive version: theaggregate.ai/model?slug=owl-alpha · How It Works · Data refreshed daily, snapshot 2026-09-19.