ConvRe — leaderboard

ConvRe evaluates language models' understanding of converse relations through Re2Text and Text2Re tasks in easy and hard settings.

Metric: Average Score (%). Source: huggingface.co. Status: saturated. 10 models tracked.

Top models

#ModelScore
1text-davinci-00365
2GPT-3.5 Turbo (0301)60.6
3Claude Instant 1.157.9
4GPT-4 (0314)56.5
5flan-t5-xl52
6flan-t5-large51.2
7flan-t5-base50.8
8flan-t5-xxl50.4
9flan-t5-small49.5

Interactive version: theaggregate.ai/benchmark?slug=convre · How the rankings work · Data refreshed daily, snapshot 2026-07-22.