An advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly improves upon its predecessor, DeepSeek-VL. DeepSeek-VL2 demonstrates superior capabilities across various tasks, including but not limited to visual question answering, optical character recognition, document/table/chart understanding, and visual grounding.
发布日期2024年12月13日
参数规模3B
上下文长度—
许可证deepseek
知识截止—
Benchmarks
评测成绩
| 评测基准 | 类别 | 分数 | 来源 |
|---|---|---|---|
| DocVQA | vision multimodal | 88.9 | 来源 |
| ChartQA | reasoning vision multimodal | 81.0 | 来源 |
| OCRBench | vision image-to-text | 80.9 | 来源 |
| TextVQA | vision multimodal image-to-text | 80.7 | 来源 |
| AI2D | vision reasoning multimodal | 71.6 | 来源 |
| MMBench | vision multimodal reasoning | 69.2 | 来源 |
| MMBench-V1.1 | vision multimodal reasoning | 68.3 | 来源 |
| InfoVQA | vision multimodal | 66.1 | 来源 |
| RealWorldQA | vision spatial_reasoning | 64.2 | 来源 |
| MathVista | math vision multimodal | 53.6 | 来源 |
| MMT-Bench | vision multimodal reasoning general | 53.2 | 来源 |
| MMStar | vision multimodal reasoning general | 45.9 | 来源 |
| MMMU | multimodal reasoning general | 40.7 | 来源 |
| MME | vision multimodal reasoning | 19.2 | 来源 |
Pricing
API 价格对比
暂无 API 价格。