Benchmark

POPE

vision safety multimodal multimodal

Polling-based Object Probing Evaluation (POPE) is a benchmark for evaluating object hallucination in Large Vision-Language Models (LVLMs). POPE addresses the problem where LVLMs generate objects inconsistent with target images by using a polling-based query method that asks yes/no questions about object presence in images, providing more stable and flexible evaluation of object hallucination.

语言EN
满分1
参评模型2

模型排名

名次 模型 机构 分数 来源
1 Phi-3.5-vision-instruct Microsoft 86.1 来源 ↗
2 Phi-4-multimodal-instruct Microsoft 85.6 来源 ↗