Ember-1: benchmark results
Provider: Other.
Unified ELO 1862 ± 32, rank #34 of 1610 rated models, from 22 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Vals AI Vibe Code Bench | 83.87 | Accuracy (%) | 86.5 |
| Vals AI Tax Agent Bench | 68.04 | Accuracy (%) | 82.8 |
| Vals AI Legal Research Bench | 44.71 | Accuracy (%) | 77.8 |
| Vals AI Harvey Legal Agent Bench | 9.17 | Accuracy (%) | 76.7 |
| Gert Labs Rankings | 54.52 | GScore (%) | 75.8 |
| Tinybird AI SQL Benchmark - Success Rate | 100 | Questions answered with a valid query within 3 retries (%) | 72.3 |
| Tinybird AI SQL Benchmark - Exactness | 51.73 | Result exactness vs human reference queries (0-100) | 71.6 |
| Vals AI Public Benefits Bench | 66.17 | Accuracy (%) | 65.2 |
| Vals AI Excel Modeling | 63.37 | Accuracy (%) | 63.8 |
| Vals AI ProofBench | 61 | Accuracy (%) | 61.2 |
| Vals AI Finance Agent v2 | 51.85 | Accuracy (%) | 57.5 |
| Vals AI Terminal-Bench 4.0 | 19.7 | Accuracy (%) | 53.6 |
Interactive version: theaggregate.ai/model?slug=ember-1 · How It Works · Data refreshed daily, snapshot 2026-10-03.