Phi-4-multimodal-instruct is a lightweight (5.57B parameters) open multimodal foundation model that leverages research and datasets from Phi-3.5 and 4.0. It processes text, image, and audio inputs to generate text outputs, supporting a 128K token context length. Enhanced via SFT, DPO, and RLHF for instruction following and safety.
发布日期2025年2月1日
参数规模5.6B
上下文长度128K
许可证MIT
知识截止2024年6月1日
Benchmarks
评测成绩
| 评测基准 | 类别 | 分数 | 来源 |
|---|---|---|---|
| ScienceQA Visual | vision reasoning multimodal | 97.5 | 来源 |
| DocVQA | vision multimodal | 93.2 | 来源 |
| MMBench | vision multimodal reasoning | 86.7 | 来源 |
| POPE | vision safety multimodal | 85.6 | 来源 |
| OCRBench | vision image-to-text | 84.4 | 来源 |
| AI2D | vision reasoning multimodal | 82.3 | 来源 |
| ChartQA | reasoning vision multimodal | 81.4 | 来源 |
| TextVQA | vision multimodal image-to-text | 75.6 | 来源 |
| InfoVQA | vision multimodal | 72.7 | 来源 |
| MathVista | math vision multimodal | 62.4 | 来源 |
| BLINK | vision multimodal reasoning | 61.3 | 来源 |
| MMMU | multimodal reasoning general | 55.1 | 来源 |
| Video-MME | multimodal vision reasoning | 55.0 | 来源 |
| InterGPS | math spatial_reasoning | 48.6 | 来源 |
| MMMU-Pro | vision multimodal reasoning general | 38.5 | 来源 |
Pricing
API 价格对比
| 服务商 | 输入价 | 输出价 | 上下文 | 吞吐(tok/s) | 延迟(s) | 函数调用 | 代码执行 | 联网搜索 |
|---|---|---|---|---|---|---|---|---|
| NanoGPT | $0.07 | $0.11 | 128K | — | — | ✗ | ✗ | ✗ |
价格单位:美元/百万 token,数据来自社区整理,仅供参考。