Benchmark

VoiceBench Avg

general reasoning safety communication multimodal

VoiceBench is the first benchmark designed to provide a multi-faceted evaluation of LLM-based voice assistants, evaluating capabilities including general knowledge, instruction-following, reasoning, and safety using both synthetic and real spoken instruction data with diverse speaker characteristics and environmental conditions.

语言EN
满分1
参评模型1

模型排名

名次 模型 机构 分数 来源
1 Qwen2.5-Omni-7B Alibaba Cloud / Qwen Team 74.1 来源 ↗