DeepSeek-V2.5 is an upgraded version that combines DeepSeek-V2-Chat and DeepSeek-Coder-V2-Instruct, integrating general and coding abilities. It better aligns with human preferences and has been optimized in various aspects, including writing and instruction following.
发布日期2024年5月8日
参数规模236B
上下文长度—
许可证deepseek
知识截止—
Benchmarks
评测成绩
| 评测基准 | 类别 | 分数 | 来源 |
|---|---|---|---|
| GSM8k | math reasoning | 95.1 | 来源 |
| MT-Bench | communication reasoning general roleplay | 90.2 | 来源 |
| HumanEval | reasoning code | 89.0 | 来源 |
| BBH | reasoning math language | 84.3 | 来源 |
| AlignBench | general language math reasoning roleplay | 80.4 | 来源 |
| MMLU | general reasoning language math | 80.4 | 来源 |
| DS-FIM-Eval | general | 78.3 | 来源 |
| Arena Hard | general reasoning creativity | 76.2 | 来源 |
| MATH | math reasoning | 74.7 | 来源 |
| HumanEval-Mul | reasoning | 73.8 | 来源 |
| Aider | reasoning code | 72.2 | 来源 |
| DS-Arena-Code | reasoning | 63.1 | 来源 |
| AlpacaEval 2.0 | general creativity reasoning | 50.5 | 来源 |
| LiveCodeBench(01-09) | reasoning general | 41.8 | 来源 |
| SWE-Bench Verified | reasoning frontend_development code | 16.8 | 来源 |
Pricing
API 价格对比
暂无 API 价格。