Qwen2.5-VL is a vision-language model from the Qwen family. Key enhancements include visual understanding (objects, text, charts, layouts), visual agent capabilities (tool use, computer/phone control), long video comprehension with event pinpointing, visual localization (bounding boxes/points), and structured output generation.
发布日期2025年1月26日
参数规模8.3B
上下文长度—
许可证Apache 2.0
知识截止—
Benchmarks
评测成绩
| 评测基准 | 类别 | 分数 | 来源 |
|---|---|---|---|
| DocVQA | vision multimodal | 95.7 | 来源 |
| MobileMiniWob++_SR | multimodal frontend_development | 91.4 | 来源 |
| Android Control Low_EM | multimodal reasoning | 91.4 | 来源 |
| ChartQA | reasoning vision multimodal | 87.3 | 来源 |
| OCRBench | vision image-to-text | 86.4 | 来源 |
| TextVQA | vision multimodal image-to-text | 84.9 | 来源 |
| ScreenSpot | vision multimodal spatial_reasoning | 84.7 | 来源 |
| MMBench | vision multimodal reasoning | 84.3 | 来源 |
| InfoVQA | vision multimodal | 82.6 | 来源 |
| AITZ_EM | multimodal reasoning | 81.9 | 来源 |
| CC-OCR | vision multimodal text-to-image | 77.8 | 来源 |
| TempCompass | vision multimodal reasoning | 71.7 | 来源 |
| VideoMME w sub. | vision multimodal video | 71.6 | 来源 |
| PerceptionTest | video multimodal reasoning physics spatial_reasoning | 70.5 | 来源 |
| MLVU | video multimodal long_context | 70.2 | 来源 |
| MVBench | vision video multimodal spatial_reasoning reasoning | 69.6 | 来源 |
| MathVista-Mini | math vision multimodal | 68.2 | 来源 |
| MMVet | vision multimodal reasoning general spatial_reasoning math | 67.1 | 来源 |
| VideoMME w/o sub. | multimodal video vision | 65.1 | 来源 |
| MMStar | vision multimodal reasoning general | 63.9 | 来源 |
| MMT-Bench | vision multimodal reasoning general | 63.6 | 来源 |
| Android Control High_EM | multimodal reasoning | 60.1 | 来源 |
| MMMU | multimodal reasoning general | 58.6 | 来源 |
| LongVideoBench | vision long_context multimodal | 54.7 | 来源 |
| Hallusion Bench | vision reasoning | 52.9 | 来源 |
| LVBench | vision multimodal long_context | 45.3 | 来源 |
| CharadesSTA | video language multimodal | 43.6 | 来源 |
| MMMU-Pro | vision multimodal reasoning general | 38.3 | 来源 |
| ScreenSpot Pro | vision multimodal spatial_reasoning | 29.0 | 来源 |
| AndroidWorld_SR | general multimodal reasoning | 25.5 | 来源 |
| MathVision | math vision multimodal | 25.1 | 来源 |
| MMBench-Video | video multimodal reasoning | 1.8 | 来源 |
Pricing
API 价格对比
暂无 API 价格。