An advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly improves upon its predecessor, DeepSeek-VL. DeepSeek-VL2 demonstrates superior capabilities across various tasks, including but not limited to visual question answering, optical character recognition, document/table/chart understanding, and visual grounding.
发布日期2024年12月13日
参数规模27B
上下文长度—
许可证deepseek
知识截止—
Benchmarks
评测成绩
| 评测基准 | 类别 | 分数 | 来源 |
|---|---|---|---|
| DocVQA | vision multimodal | 93.3 | 来源 |
| ChartQA | reasoning vision multimodal | 86.0 | 来源 |
| TextVQA | vision multimodal image-to-text | 84.2 | 来源 |
| AI2D | vision reasoning multimodal | 81.4 | 来源 |
| OCRBench | vision image-to-text | 81.1 | 来源 |
| MMBench | vision multimodal reasoning | 79.6 | 来源 |
| MMBench-V1.1 | vision multimodal reasoning | 79.2 | 来源 |
| InfoVQA | vision multimodal | 78.1 | 来源 |
| RealWorldQA | vision spatial_reasoning | 68.4 | 来源 |
| MMT-Bench | vision multimodal reasoning general | 63.6 | 来源 |
| MathVista | math vision multimodal | 62.8 | 来源 |
| MMStar | vision multimodal reasoning general | 61.3 | 来源 |
| MMMU | multimodal reasoning general | 51.1 | 来源 |
| MME | vision multimodal reasoning | 22.5 | 来源 |
Pricing
API 价格对比
暂无 API 价格。