Phi-3.5-vision-instruct is a 4.2B-parameter open multimodal model with up to 128K context tokens. It emphasizes multi-frame image understanding and reasoning, boosting performance on single-image benchmarks while enabling multi-image comparison, summarization, and even video analysis. The model underwent safety post-training for improved instruction-following, alignment, and robust handling of visual and text inputs, and is released under the MIT license.
发布日期2024年8月23日
参数规模4.2B
上下文长度—
许可证MIT
知识截止—
Benchmarks
评测成绩
| 评测基准 | 类别 | 分数 | 来源 |
|---|---|---|---|
| ScienceQA | reasoning math multimodal | 91.3 | 来源 |
| POPE | vision safety multimodal | 86.1 | 来源 |
| MMBench | vision multimodal reasoning | 81.9 | 来源 |
| ChartQA | reasoning vision multimodal | 81.8 | 来源 |
| AI2D | vision reasoning multimodal | 78.1 | 来源 |
| TextVQA | vision multimodal image-to-text | 72.0 | 来源 |
| MathVista | math vision multimodal | 43.9 | 来源 |
| MMMU | multimodal reasoning general | 43.0 | 来源 |
| InterGPS | math spatial_reasoning | 36.3 | 来源 |
Pricing
API 价格对比
暂无 API 价格。