DeepSeek-V3.1 is a hybrid model supporting both thinking and non-thinking modes through different chat templates. Built on DeepSeek-V3.1-Base with a two-phase long context extension (32K phase: 630B tokens, 128K phase: 209B tokens), it features 671B total parameters with 37B activated. Key improvements include smarter tool calling through post-training optimization, higher thinking efficiency achieving comparable quality to DeepSeek-R1-0528 while responding more quickly, and UE8M0 FP8 scale data format for model weights and activations. The model excels in both reasoning tasks (thinking mode) and practical applications (non-thinking mode), with particularly strong performance in code agent tasks, math competitions, and search-based problem solving.
Benchmarks
评测成绩
| 评测基准 | 类别 | 分数 | 来源 |
|---|---|---|---|
| SimpleQA | general reasoning | 93.4 | 来源 |
| MMLU-Redux | language reasoning math general | 91.8 | 来源 |
| MMLU-Pro | language reasoning math general | 83.7 | 来源 |
| GPQA | reasoning general | 74.9 | 来源 |
| Codeforces | math reasoning | 69.7 | 来源 |
| Aider-Polyglot | general code | 68.4 | 来源 |
| AIME 2024 | math reasoning | 66.3 | 来源 |
| SWE-Bench Verified | reasoning frontend_development code | 66.0 | 来源 |
| LiveCodeBench | reasoning general code | 56.4 | 来源 |
| SWE-Bench Multilingual | reasoning code | 54.5 | 来源 |
| AIME 2025 | math reasoning | 49.8 | 来源 |
| BrowseComp-zh | reasoning search | 49.2 | 来源 |
| HMMT 2025 | math | 33.5 | 来源 |
| Terminal-Bench | reasoning code | 31.3 | 来源 |
| BrowseComp | reasoning search | 30.0 | 来源 |
| Humanity's Last Exam | general | 15.9 | 来源 |
Pricing
API 价格对比
| 服务商 | 输入价 | 输出价 | 上下文 | 吞吐(tok/s) | 延迟(s) | 函数调用 | 代码执行 | 联网搜索 |
|---|---|---|---|---|---|---|---|---|
| Azure | $0.56 | $1.68 | 131K | — | — | ✓ | ✗ | ✗ |
| Azure Cognitive Services | $0.56 | $1.68 | 131K | — | — | ✓ | ✗ | ✗ |
价格单位:美元/百万 token,数据来自社区整理,仅供参考。