Benchmark

VocalSound

audio audio

A dataset for improving human vocal sounds recognition, containing over 21,000 crowdsourced recordings of laughter, sighs, coughs, throat clearing, sneezes, and sniffs from 3,365 unique subjects. Used for audio event classification and recognition of human non-speech vocalizations.

语言EN
满分1
参评模型1

模型排名

名次 模型 机构 分数 来源
1 Qwen2.5-Omni-7B Alibaba Cloud / Qwen Team 93.9 来源 ↗