Benchmark
MT-Bench
communication
reasoning
general
roleplay
text
MT-Bench is a challenging multi-turn benchmark that measures the ability of large language models to engage in coherent, informative, and engaging conversations. It uses strong LLMs as judges for scalable and explainable evaluation of multi-turn dialogue capabilities.
语言EN
满分100
参评模型11
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Qwen2.5 72B Instruct | Alibaba Cloud / Qwen Team | 93.5 | 来源 ↗ |
| 2 | Llama-3.3 Nemotron Super 49B v1 | NVIDIA | 91.7 | 来源 ↗ |
| 3 | DeepSeek-V2.5 | DeepSeek | 90.2 | 来源 ↗ |
| 4 | Qwen2.5 7B Instruct | Alibaba Cloud / Qwen Team | 87.5 | 来源 ↗ |
| 5 | Mistral Large 2 | Mistral AI | 86.3 | 来源 ↗ |
| 6 | Qwen2 7B Instruct | Alibaba Cloud / Qwen Team | 84.1 | 来源 ↗ |
| 7 | Mistral Small 3 24B Instruct | Mistral AI | 83.5 | 来源 ↗ |
| 8 | Ministral 8B Instruct | Mistral AI | 83.0 | 来源 ↗ |
| 9 | Llama 3.1 Nemotron Nano 8B V1 | NVIDIA | 81.0 | 来源 ↗ |
| 10 | Pixtral-12B | Mistral AI | 76.8 | 来源 ↗ |
| 11 | Llama 3.1 Nemotron 70B Instruct | NVIDIA | 9.0 | 来源 ↗ |