Model

DeepSeek-V3 0324

DeepSeek
开源 MIT + Model License (Commercial use allowed)

A powerful Mixture-of-Experts (MoE) language model with 671B total parameters (37B activated per token). Features Multi-head Latent Attention (MLA), auxiliary-loss-free load balancing, and multi-token prediction training. Pre-trained on 14.8T tokens with strong performance in reasoning, math, and code tasks.

发布日期2025年3月25日
参数规模671B
上下文长度131K
许可证MIT + Model License (Commercial use allowed)
知识截止

Benchmarks

评测成绩

评测基准 类别 分数 来源
MATH-500 math reasoning 94.0 来源
MMLU-Pro language reasoning math general 81.2 来源
GPQA reasoning general 68.4 来源
AIME 2024 math reasoning 59.4 来源
LiveCodeBench reasoning general code 49.2 来源

Pricing

API 价格对比

服务商 输入价 输出价 上下文 吞吐(tok/s) 延迟(s) 函数调用 代码执行 联网搜索
NanoGPT $0.25 $0.70 128K
Cortecs $0.55 $1.65 128K
Azure $1.14 $4.56 131K
Azure Cognitive Services $1.14 $4.56 131K

价格单位:美元/百万 token,数据来自社区整理,仅供参考。