SpacetimeDB LLM Benchmark (TypeScript) — leaderboard
TypeScript subset of the SpacetimeDB LLM Benchmark. Tests LLM ability to generate correct SpacetimeDB TypeScript module code, evaluating basics (reducers, tables, CRUD) and schema patterns against live module execution.
Metric: Task Pass Rate (%). Source: spacetimedb.com. Status: saturation imminent. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.6 | 89.4 |
| 2 | Claude Opus 4.6 | 89.4 |
| 3 | Gemini 3.1 Pro (Preview) | 84.1 |
| 4 | Gemini 3 Flash | 81.8 |
| 5 | DeepSeek V3 Chat | 77.6 |
| 6 | GPT-5 Mini | 66.1 |
| 7 | Grok 4 | 65.9 |
| 8 | GPT-5.2 Codex | 62.1 |
| 9 | DeepSeek Reasoner | 50.8 |
Interactive version: theaggregate.ai/benchmark?slug=spacetimedb-llm-benchmark-typescript · How the rankings work · Data refreshed daily, snapshot 2026-07-22.