Qwen2.5-Coder is a specialized coding model trained on 5.5 trillion tokens of code data, supporting 92 programming languages with a 128K context window. It excels in code generation, completion, and repair while maintaining strong performance in math and general tasks. The model demonstrates exceptional capabilities in multi-programming language tasks and code reasoning.
发布日期2024年9月19日
参数规模7B
上下文长度—
许可证Apache 2.0
知识截止—
Benchmarks
评测成绩
| 评测基准 | 类别 | 分数 | 来源 |
|---|---|---|---|
| HumanEval | reasoning code | 88.4 | 来源 |
| GSM8k | math reasoning | 83.9 | 来源 |
| MBPP | reasoning general | 83.5 | 来源 |
| HellaSwag | reasoning | 76.8 | 来源 |
| Winogrande | reasoning language | 72.9 | 来源 |
| MMLU-Base | language reasoning math general | 68.0 | 来源 |
| MMLU | general reasoning language math | 67.6 | 来源 |
| MMLU-Redux | language reasoning math general | 66.6 | 来源 |
| ARC-C | reasoning general | 60.9 | 来源 |
| CRUXEval-Input-CoT | reasoning | 56.5 | 来源 |
| CRUXEval-Output-CoT | reasoning | 56.0 | 来源 |
| Aider | reasoning code | 55.6 | 来源 |
| TruthfulQA | general reasoning legal healthcare finance | 50.6 | 来源 |
| MATH | math reasoning | 46.6 | 来源 |
| BigCodeBench | general reasoning | 41.0 | 来源 |
| MMLU-Pro | language reasoning math general | 40.1 | 来源 |
| STEM | math reasoning multimodal | 34.0 | 来源 |
| TheoremQA | math reasoning physics finance | 34.0 | 来源 |
| LiveCodeBench | reasoning general code | 18.2 | 来源 |
Pricing
API 价格对比
暂无 API 价格。