Benchmark
MTVQA
vision
multimodal
text-to-image
multimodal
多语言
MTVQA (Multilingual Text-Centric Visual Question Answering) is the first benchmark featuring high-quality human expert annotations across 9 diverse languages, consisting of 6,778 question-answer pairs across 2,116 images. It addresses visual-textual misalignment problems in multilingual text-centric VQA.
语言EN
满分1
参评模型1
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Qwen2-VL-72B-Instruct | Alibaba Cloud / Qwen Team | 30.9 | 来源 ↗ |