Ember-1: benchmark results

Provider: Other.

Unified ELO 1862 ± 32, rank #34 of 1610 rated models, from 22 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Vals AI Vibe Code Bench83.87Accuracy (%)86.5
Vals AI Tax Agent Bench68.04Accuracy (%)82.8
Vals AI Legal Research Bench44.71Accuracy (%)77.8
Vals AI Harvey Legal Agent Bench9.17Accuracy (%)76.7
Gert Labs Rankings54.52GScore (%)75.8
Tinybird AI SQL Benchmark - Success Rate100Questions answered with a valid query within 3 retries (%)72.3
Tinybird AI SQL Benchmark - Exactness51.73Result exactness vs human reference queries (0-100)71.6
Vals AI Public Benefits Bench66.17Accuracy (%)65.2
Vals AI Excel Modeling63.37Accuracy (%)63.8
Vals AI ProofBench61Accuracy (%)61.2
Vals AI Finance Agent v251.85Accuracy (%)57.5
Vals AI Terminal-Bench 4.019.7Accuracy (%)53.6

Interactive version: theaggregate.ai/model?slug=ember-1 · How It Works · Data refreshed daily, snapshot 2026-10-03.