SpacetimeDB LLM Benchmark (C#) — leaderboard

C# subset of the SpacetimeDB LLM Benchmark. Tests LLM ability to generate correct SpacetimeDB C# module code, evaluating basics (reducers, tables, CRUD) and schema patterns against live module execution.

Metric: Task Pass Rate (%). Source: spacetimedb.com. 10 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.696.2
2Grok 494.7
3DeepSeek V3 Chat88.9
4Claude Opus 4.687.9
5Gemini 3.1 Pro (Preview)85.6
6GPT-5 Mini81.1
7Gemini 3 Flash79.5
8DeepSeek Reasoner72.7
9GPT-5.2 Codex68

Interactive version: theaggregate.ai/benchmark?slug=spacetimedb-llm-benchmark-c · How the rankings work · Data refreshed daily, snapshot 2026-07-22.