Benchmark

PathMCQA

healthcare vision multimodal reasoning multimodal

PathMMU is a massive multimodal expert-level benchmark for understanding and reasoning in pathology, containing 33,428 multimodal multi-choice questions and 24,067 images validated by seven pathologists. It evaluates Large Multimodal Models (LMMs) performance on pathology tasks, with the top-performing model GPT-4V achieving only 49.8% zero-shot performance compared to 71.8% for human pathologists.

语言EN
满分1
参评模型1

模型排名

名次 模型 机构 分数 来源
1 MedGemma 4B IT Google 69.8 来源 ↗