SpacetimeDB LLM Benchmark (Rust): leaderboard
Rust subset of the SpacetimeDB LLM Benchmark. Tests LLM ability to generate correct SpacetimeDB Rust module code, evaluating basics (reducers, tables, CRUD) and schema patterns against live module execution.
Metric: Eval Pass Rate (%). Source: spacetimedb.com. Status: saturation imminent. 14 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Grok 4.3 | 97.8 |
| 2 | Grok Build 0.1 | 97.8 |
| 3 | Claude Sonnet 4.6 | 96.6 |
| 4 | GPT-5.5 | 96.6 |
| 5 | Claude Opus 4.8 | 94.4 |
| 6 | GPT-5.4 Mini | 92.1 |
| 7 | GPT-5.6 Sol | 82.9 |
| 8 | DeepSeek V4 Pro | 68.3 |
| 9 | Claude Fable 5 | 68.3 |
| 10 | Claude Opus 5 | 68.3 |
| 11 | Gemini 3.1 Pro (Preview) | 67.5 |
| 12 | Kimi K3 | 65.1 |
| 13 | DeepSeek V4 Flash | 62.7 |
| 14 | Gemini 3.5 Flash | 49.4 |
Interactive version: theaggregate.ai/benchmark?slug=spacetimedb-llm-benchmark-rust · How It Works · Data refreshed daily, snapshot 2026-09-05.