Benchmark

MT-Bench

communication reasoning general roleplay text

MT-Bench is a challenging multi-turn benchmark that measures the ability of large language models to engage in coherent, informative, and engaging conversations. It uses strong LLMs as judges for scalable and explainable evaluation of multi-turn dialogue capabilities.

语言EN
满分100
参评模型11

模型排名

名次 模型 机构 分数 来源
1 Qwen2.5 72B Instruct Alibaba Cloud / Qwen Team 93.5 来源 ↗
2 Llama-3.3 Nemotron Super 49B v1 NVIDIA 91.7 来源 ↗
3 DeepSeek-V2.5 DeepSeek 90.2 来源 ↗
4 Qwen2.5 7B Instruct Alibaba Cloud / Qwen Team 87.5 来源 ↗
5 Mistral Large 2 Mistral AI 86.3 来源 ↗
6 Qwen2 7B Instruct Alibaba Cloud / Qwen Team 84.1 来源 ↗
7 Mistral Small 3 24B Instruct Mistral AI 83.5 来源 ↗
8 Ministral 8B Instruct Mistral AI 83.0 来源 ↗
9 Llama 3.1 Nemotron Nano 8B V1 NVIDIA 81.0 来源 ↗
10 Pixtral-12B Mistral AI 76.8 来源 ↗
11 Llama 3.1 Nemotron 70B Instruct NVIDIA 9.0 来源 ↗