A multimodal model capable of processing text and visual information, including documents, diagrams, charts, screenshots, and photographs. Notable for strong real-world spatial understanding capabilities.
发布日期2024年4月12日
参数规模—
上下文长度—
许可证Proprietary
知识截止—
Benchmarks
评测成绩
| 评测基准 | 类别 | 分数 | 来源 |
|---|---|---|---|
| AI2D | vision reasoning multimodal | 88.3 | 来源 |
| DocVQA | vision multimodal | 85.6 | 来源 |
| TextVQA | vision multimodal image-to-text | 78.1 | 来源 |
| ChartQA | reasoning vision multimodal | 76.1 | 来源 |
| RealWorldQA | vision spatial_reasoning | 68.7 | 来源 |
| MMMU | multimodal reasoning general | 53.6 | 来源 |
| MathVista | math vision multimodal | 52.8 | 来源 |
Pricing
API 价格对比
暂无 API 价格。