A 12B parameter multimodal model with a 400M parameter vision encoder, capable of understanding both natural images and documents. Excels at multimodal tasks while maintaining strong text-only performance. Supports variable image sizes and multiple images in context.
发布日期2024年9月17日
参数规模12.4B
上下文长度128K
许可证Apache 2.0
知识截止—
Benchmarks
评测成绩
| 评测基准 | 类别 | 分数 | 来源 |
|---|---|---|---|
| DocVQA | vision multimodal | 90.7 | 来源 |
| ChartQA | reasoning vision multimodal | 81.8 | 来源 |
| VQAv2 | vision multimodal reasoning | 78.6 | 来源 |
| MT-Bench | communication reasoning general roleplay | 76.8 | 来源 |
| HumanEval | reasoning code | 72.0 | 来源 |
| MMLU | general reasoning language math | 69.2 | 来源 |
| IFEval | general | 61.3 | 来源 |
| MM-MT-Bench | multimodal communication | 60.5 | 来源 |
| MathVista | math vision multimodal | 58.0 | 来源 |
| MM IF-Eval | multimodal reasoning | 52.7 | 来源 |
| MMMU | multimodal reasoning general | 52.5 | 来源 |
| MATH | math reasoning | 48.1 | 来源 |
Pricing
API 价格对比
| 服务商 | 输入价 | 输出价 | 上下文 | 吞吐(tok/s) | 延迟(s) | 函数调用 | 代码执行 | 联网搜索 |
|---|---|---|---|---|---|---|---|---|
| Mistral | $0.15 | $0.15 | 128K | — | — | ✓ | ✗ | ✗ |
| Scaleway | $0.20 | $0.20 | 128K | — | — | ✓ | ✗ | ✗ |
价格单位:美元/百万 token,数据来自社区整理,仅供参考。