Benchmark

MMAU Speech

audio multimodal reasoning speech-to-text multimodal

A subset of the MMAU benchmark focused specifically on speech understanding and reasoning tasks. Part of a comprehensive multimodal audio understanding benchmark that evaluates models on expert-level knowledge and complex reasoning across speech audio clips.

语言EN
满分1
参评模型1

模型排名

名次 模型 机构 分数 来源
1 Qwen2.5-Omni-7B Alibaba Cloud / Qwen Team 59.8 来源 ↗