A multimodal model capable of processing audio, images, video, and text with high efficiency. Features JSON mode, function calling, code execution, and system instructions support. Optimized for fast inference with 8B parameters.
发布日期2024年3月15日
参数规模8B
上下文长度—
许可证Proprietary
知识截止2024年10月1日
Benchmarks
评测成绩
| 评测基准 | 类别 | 分数 | 来源 |
|---|---|---|---|
| XSTest | safety | 92.6 | 来源 |
| FLEURS | language speech-to-text | 86.4 | 来源 |
| Natural2Code | reasoning general | 75.5 | 来源 |
| WMT23 | language | 72.6 | 来源 |
| Video-MME | multimodal vision reasoning | 66.2 | 来源 |
| MATH | math reasoning | 58.7 | 来源 |
| MMLU-Pro | language reasoning math general | 58.7 | 来源 |
| MathVista | math vision multimodal | 54.7 | 来源 |
| MRCR | long_context reasoning general | 54.7 | 来源 |
| MMMU | multimodal reasoning general | 53.7 | 来源 |
| Vibe-Eval | multimodal vision general | 40.9 | 来源 |
| GPQA | reasoning general | 38.4 | 来源 |
| HiddenMath | math reasoning | 32.8 | 来源 |
Pricing
API 价格对比
暂无 API 价格。