Llama 4 Scout is a natively multimodal model capable of processing both text and images. It features a 17 billion activated parameter (109B total) mixture-of-experts (MoE) architecture with 16 experts, supporting a wide range of multimodal tasks such as conversational interaction, image analysis, and code generation. The model includes a 10 million token context window.
发布日期2025年4月5日
参数规模109B
上下文长度131K
许可证Llama 4 Community License Agreement
知识截止—
Benchmarks
评测成绩
| 评测基准 | 类别 | 分数 | 来源 |
|---|---|---|---|
| DocVQA | vision multimodal | 94.4 | 来源 |
| MGSM | math reasoning | 90.6 | 来源 |
| ChartQA | reasoning vision multimodal | 88.8 | 来源 |
| MMLU | general reasoning language math | 79.6 | 来源 |
| MMLU-Pro | language reasoning math general | 74.3 | 来源 |
| MathVista | math vision multimodal | 70.7 | 来源 |
| MMMU | multimodal reasoning general | 69.4 | 来源 |
| MBPP | reasoning general | 67.8 | 来源 |
| GPQA | reasoning general | 57.2 | 来源 |
| MATH | math reasoning | 50.3 | 来源 |
| LiveCodeBench | reasoning general code | 32.8 | 来源 |
| TydiQA | language reasoning | 31.5 | 来源 |
Pricing
API 价格对比
| 服务商 | 输入价 | 输出价 | 上下文 | 吞吐(tok/s) | 延迟(s) | 函数调用 | 代码执行 | 联网搜索 |
|---|---|---|---|---|---|---|---|---|
| Helicone | $0.08 | $0.30 | 131K | — | — | ✓ | ✗ | ✗ |
价格单位:美元/百万 token,数据来自社区整理,仅供参考。