An instruction-tuned, large multimodal model that excels at visual understanding and step-by-step reasoning. It supports image and video input, with dynamic resolution handling and improved positional embeddings (M-ROPE), enabling advanced capabilities such as complex problem solving, multilingual text recognition in images, and agent-like interactions in video contexts.
发布日期2024年8月29日
参数规模73.4B
上下文长度—
许可证tongyi-qianwen
知识截止2023年6月30日
Benchmarks
评测成绩
| 评测基准 | 类别 | 分数 | 来源 |
|---|---|---|---|
| DocVQAtest | vision multimodal | 96.5 | 来源 |
| VCR_en_easy | vision reasoning | 91.9 | 来源 |
| ChartQA | reasoning vision multimodal | 88.3 | 来源 |
| OCRBench | vision image-to-text | 87.7 | 来源 |
| MMBench_test | vision multimodal reasoning | 86.5 | 来源 |
| TextVQA | vision multimodal image-to-text | 85.5 | 来源 |
| InfoVQAtest | vision multimodal | 84.5 | 来源 |
| EgoSchema | vision reasoning long_context | 77.9 | 来源 |
| RealWorldQA | vision spatial_reasoning | 77.8 | 来源 |
| MMVetGPT4Turbo | vision multimodal reasoning general spatial_reasoning math | 74.0 | 来源 |
| MVBench | vision video multimodal spatial_reasoning reasoning | 73.6 | 来源 |
| MathVista-Mini | math vision multimodal | 70.5 | 来源 |
| MMMUval | vision general reasoning multimodal | 64.5 | 来源 |
| MMMU-Pro | vision multimodal reasoning general | 46.2 | 来源 |
| MTVQA | vision multimodal text-to-image | 30.9 | 来源 |
Pricing
API 价格对比
暂无 API 价格。