VNTL Leaderboard — leaderboard

Japanese visual-novel translation benchmark where LLMs and translation systems render 256 samples into English while preserving semantic similarity.

Metric: Accuracy (%). Source: huggingface.co. Status: saturation imminent. 87 models tracked.

Top models

#ModelScore
1GPT-4o (2024-05-13)75.16
2GPT-4o (2024-08-06)74.97
3Claude 3 Opus74.59
4Claude 3.5 Sonnet (20240620)74.4
5DeepSeek V3 Chat74.24
6Claude 3.5 Sonnet (20241022)72.8
7GPT-4o Mini (2024-07-18)72.23
8Grok 2 (1212)71.6
9Grok Beta71.27
10DeepSeek V2.571.14
11Qwen 2 72B Instruct70.2
12GPT-3.5 Turbo (1106)69.98
13Llama 3.1 70B Instruct69.79
14Llama 3.1 405B Instruct69.46
15GPT-4 (0613)69.28

Interactive version: theaggregate.ai/benchmark?slug=vntl-leaderboard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.