A powerful Mixture-of-Experts (MoE) language model with 671B total parameters (37B activated per token). Features Multi-head Latent Attention (MLA), auxiliary-loss-free load balancing, and multi-token prediction training. Pre-trained on 14.8T tokens with strong performance in reasoning, math, and code tasks.
发布日期2025年3月25日
参数规模671B
上下文长度131K
许可证MIT + Model License (Commercial use allowed)
知识截止—
Benchmarks
评测成绩
Pricing
API 价格对比
| 服务商 | 输入价 | 输出价 | 上下文 | 吞吐(tok/s) | 延迟(s) | 函数调用 | 代码执行 | 联网搜索 |
|---|---|---|---|---|---|---|---|---|
| NanoGPT | $0.25 | $0.70 | 128K | — | — | ✓ | ✗ | ✗ |
| Cortecs | $0.55 | $1.65 | 128K | — | — | ✓ | ✗ | ✗ |
| Azure | $1.14 | $4.56 | 131K | — | — | ✓ | ✗ | ✗ |
| Azure Cognitive Services | $1.14 | $4.56 | 131K | — | — | ✓ | ✗ | ✗ |
价格单位:美元/百万 token,数据来自社区整理,仅供参考。