TemporalBench: leaderboard

TemporalBench evaluates LLM-based agents on contextual and event-informed time-series tasks spanning multiple datasets and task types.

Metric: Overall MCQ Accuracy (%). Source: huggingface.co. 13 models tracked.

Top models

#ModelScore
1TimeClaw (deepseek-v4-pro)41.8
2Single LLM (deepseek-chat)36.47
3AgentScope (deepseek-chat)35.21
4CAMEL (deepseek-chat)35.11
5MetaGPT (deepseek-chat)34.42
6Single LLM (gpt-4o)34.11
7AgentScope (gpt-4o)32.91
8TimeCopilot (deepseek-chat)32.34
9CAMEL (gpt-4o)32.33
10MetaGPT (gpt-4o)31.56

Interactive version: theaggregate.ai/benchmark?slug=temporalbench · How It Works · Data refreshed daily, snapshot 2026-09-05.