Benchmark

MM-MT-Bench

multimodal communication multimodal

A multi-turn LLM-as-a-judge evaluation benchmark for testing multimodal instruction-tuned models' ability to follow user instructions in multi-turn dialogues and answer open-ended questions in a zero-shot manner.

语言EN
满分100
参评模型3

模型排名

名次 模型 机构 分数 来源
1 Pixtral Large Mistral AI 74.0 来源 ↗
2 Pixtral-12B Mistral AI 60.5 来源 ↗
3 Qwen2.5-Omni-7B Alibaba Cloud / Qwen Team 6.0 来源 ↗