An advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly improves upon its predecessor, DeepSeek-VL. DeepSeek-VL2 demonstrates superior capabilities across various tasks, including but not limited to visual question answering, optical character recognition, document/table/chart understanding, and visual grounding.
发布日期2024年12月13日
参数规模16B
上下文长度—
许可证deepseek
知识截止—
Benchmarks
评测成绩
| 评测基准 | 类别 | 分数 | 来源 |
|---|---|---|---|
| DocVQA | vision multimodal | 92.3 | 来源 |
| ChartQA | reasoning vision multimodal | 84.5 | 来源 |
| OCRBench | vision image-to-text | 83.4 | 来源 |
| TextVQA | vision multimodal image-to-text | 83.4 | 来源 |
| MMBench | vision multimodal reasoning | 80.3 | 来源 |
| AI2D | vision reasoning multimodal | 80.0 | 来源 |
| MMBench-V1.1 | vision multimodal reasoning | 79.3 | 来源 |
| InfoVQA | vision multimodal | 75.8 | 来源 |
| RealWorldQA | vision spatial_reasoning | 65.4 | 来源 |
| MMT-Bench | vision multimodal reasoning general | 62.9 | 来源 |
| MathVista | math vision multimodal | 60.7 | 来源 |
| MMStar | vision multimodal reasoning general | 57.0 | 来源 |
| MMMU | multimodal reasoning general | 48.0 | 来源 |
| MME | vision multimodal reasoning | 21.2 | 来源 |
Pricing
API 价格对比
暂无 API 价格。