Benchmark
DocVQAtest
vision
multimodal
multimodal
DocVQA is a Visual Question Answering benchmark on document images containing 50,000 questions defined on 12,000+ document images. The benchmark focuses on understanding document structure and content to answer questions about various document types including letters, memos, notes, and reports from the UCSF Industry Documents Library.
语言EN
满分1
参评模型1
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Qwen2-VL-72B-Instruct | Alibaba Cloud / Qwen Team | 96.5 | 来源 ↗ |