SpacetimeDB LLM Benchmark (TypeScript): leaderboard
TypeScript subset of the SpacetimeDB LLM Benchmark. Tests LLM ability to generate correct SpacetimeDB TypeScript module code, evaluating basics (reducers, tables, CRUD) and schema patterns against live module execution.
Metric: Eval Pass Rate (%). Source: spacetimedb.com. Status: saturation imminent. 14 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.6 Sol | 99.5 |
| 2 | Claude Sonnet 4.6 | 97.8 |
| 3 | GPT-5.5 | 97.8 |
| 4 | Claude Opus 4.8 | 97.8 |
| 5 | Grok 4.3 | 97.8 |
| 6 | Grok Build 0.1 | 97.8 |
| 7 | Gemini 3.1 Pro (Preview) | 97.4 |
| 8 | GPT-5.4 Mini | 96.6 |
| 9 | Claude Opus 5 | 96.4 |
| 10 | Claude Fable 5 | 95.3 |
| 11 | Kimi K3 | 93.3 |
| 12 | DeepSeek V4 Flash | 93.2 |
| 13 | DeepSeek V4 Pro | 89.1 |
| 14 | Gemini 3.5 Flash | 44.9 |
Interactive version: theaggregate.ai/benchmark?slug=spacetimedb-llm-benchmark-typescript · How It Works · Data refreshed daily, snapshot 2026-09-05.