ConvRe — leaderboard
ConvRe evaluates language models' understanding of converse relations through Re2Text and Text2Re tasks in easy and hard settings.
Metric: Average Score (%). Source: huggingface.co. Status: saturated. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | text-davinci-003 | 65 |
| 2 | GPT-3.5 Turbo (0301) | 60.6 |
| 3 | Claude Instant 1.1 | 57.9 |
| 4 | GPT-4 (0314) | 56.5 |
| 5 | flan-t5-xl | 52 |
| 6 | flan-t5-large | 51.2 |
| 7 | flan-t5-base | 50.8 |
| 8 | flan-t5-xxl | 50.4 |
| 9 | flan-t5-small | 49.5 |
Interactive version: theaggregate.ai/benchmark?slug=convre · How the rankings work · Data refreshed daily, snapshot 2026-07-22.