Benchmark
MMAU Speech
audio
multimodal
reasoning
speech-to-text
multimodal
A subset of the MMAU benchmark focused specifically on speech understanding and reasoning tasks. Part of a comprehensive multimodal audio understanding benchmark that evaluates models on expert-level knowledge and complex reasoning across speech audio clips.
语言EN
满分1
参评模型1
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Qwen2.5-Omni-7B | Alibaba Cloud / Qwen Team | 59.8 | 来源 ↗ |