SpacetimeDB LLM Benchmark (Rust) — leaderboard
Rust subset of the SpacetimeDB LLM Benchmark. Tests LLM ability to generate correct SpacetimeDB Rust module code, evaluating basics (reducers, tables, CRUD) and schema patterns against live module execution.
Metric: Task Pass Rate (%). Source: spacetimedb.com. Status: saturated. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.6 | 100 |
| 2 | Claude Opus 4.6 | 100 |
| 3 | Gemini 3 Flash | 97.7 |
| 4 | GPT-5 Mini | 94.7 |
| 5 | Grok 4 | 90.9 |
| 6 | GPT-5.2 Codex | 90.9 |
| 7 | DeepSeek V3 Chat | 83.6 |
| 8 | Gemini 3.1 Pro (Preview) | 77.3 |
| 9 | DeepSeek Reasoner | 53.3 |
Interactive version: theaggregate.ai/benchmark?slug=spacetimedb-llm-benchmark-rust · How the rankings work · Data refreshed daily, snapshot 2026-07-22.