TemporalBench - T4: leaderboard

Metric: Accuracy (%). Source: huggingface.co. 13 models tracked.

Top models

#ModelScore
1TimeClaw (deepseek-v4-pro)43.63
2AgentScope (deepseek-chat)35.08
3CAMEL (deepseek-chat)35.08
4MetaGPT (deepseek-chat)34.56
5Single LLM (deepseek-chat)28.79
6TimeCopilot (deepseek-chat)28.35
7CAMEL (gpt-4o)27.92
8AgentScope (gpt-4o)27.57
9Single LLM (gpt-4o)26.7
10TimeSeries Scientist (deepseek-chat)26.7

Interactive version: theaggregate.ai/benchmark?slug=temporalbench-t4 · How It Works · Data refreshed daily, snapshot 2026-09-05.