Model

Phi-3.5-vision-instruct

Microsoft
开源 MIT 多模态

Phi-3.5-vision-instruct is a 4.2B-parameter open multimodal model with up to 128K context tokens. It emphasizes multi-frame image understanding and reasoning, boosting performance on single-image benchmarks while enabling multi-image comparison, summarization, and even video analysis. The model underwent safety post-training for improved instruction-following, alignment, and robust handling of visual and text inputs, and is released under the MIT license.

发布日期2024年8月23日
参数规模4.2B
上下文长度
许可证MIT
知识截止

Benchmarks

评测成绩

评测基准 类别 分数 来源
ScienceQA reasoning math multimodal 91.3 来源
POPE vision safety multimodal 86.1 来源
MMBench vision multimodal reasoning 81.9 来源
ChartQA reasoning vision multimodal 81.8 来源
AI2D vision reasoning multimodal 78.1 来源
TextVQA vision multimodal image-to-text 72.0 来源
MathVista math vision multimodal 43.9 来源
MMMU multimodal reasoning general 43.0 来源
InterGPS math spatial_reasoning 36.3 来源

Pricing

API 价格对比

暂无 API 价格。