Benchmark

SlakeVQA

vision healthcare multimodal reasoning multimodal 多语言

A semantically-labeled knowledge-enhanced dataset for medical visual question answering. Contains 642 radiology images (CT scans, MRI scans, X-rays) covering five body parts and 14,028 bilingual English-Chinese question-answer pairs annotated by experienced physicians. Features comprehensive semantic labels and a structural medical knowledge base with both vision-only and knowledge-based questions requiring external medical knowledge reasoning.

语言EN
满分1
参评模型1

模型排名

名次 模型 机构 分数 来源
1 MedGemma 4B IT Google 62.3 来源 ↗