Kimi K2-Instruct-0905 is the latest, most capable version of Kimi K2, achieving state-of-the-art performance in frontier knowledge, math, and coding among non-thinking models. This Mixture-of-Experts model features 32 billion activated parameters and 1 trillion total parameters, meticulously optimized for agentic tasks. Key features include enhanced agentic coding intelligence, extended context length to 256K tokens, and a hybrid architecture trained with MuonClip optimizer on 15.5T tokens. The model achieves 65.8% on SWE-bench Verified (single attempt), 47.3% on SWE-bench Multilingual, and excels at tool use with 70.6% on Tau2-retail. It is a reflex-grade model without long thinking, designed to act and execute complex tasks seamlessly.
Benchmarks
评测成绩
| 评测基准 | 类别 | 分数 | 来源 |
|---|---|---|---|
| Math 500 | math reasoning | 97.4 | 来源 |
| Mmlu Redux | language reasoning math general | 92.7 | 来源 |
| Ifeval | general | 89.8 | 来源 |
| Mmlu | general reasoning language math | 89.5 | 来源 |
| Autologi | reasoning | 89.5 | 来源 |
| Zebralogic | reasoning | 89.0 | 来源 |
| Multiple | general language | 85.7 | 来源 |
| Mmlu Pro | language reasoning math general | 81.1 | 来源 |
| Acebench | general reasoning | 76.5 | 来源 |
| Livebench | math reasoning general | 76.4 | 来源 |
| Gpqa | reasoning general | 75.1 | 来源 |
| Cnmo 2024 | math | 74.3 | 来源 |
| Tau2 Retail | communication reasoning | 70.6 | 来源 |
| Aime 2024 | math reasoning | 69.6 | 来源 |
| Swe Bench Verified | reasoning frontend_development code | 65.8 | 来源 |
| Tau2 Telecom | communication reasoning | 65.8 | 来源 |
| Polymath En | math reasoning | 65.1 | 来源 |
| Aider Polyglot | general code | 60.0 | 来源 |
| Supergpqa | reasoning general math legal healthcare finance chemistry economics physics | 57.2 | 来源 |
| Tau2 Airline | reasoning communication | 56.5 | 来源 |
| Multichallenge | communication reasoning | 54.1 | 来源 |
| Livecodebench | reasoning general code | 53.7 | 来源 |
| Aime 2025 | math reasoning | 49.5 | 来源 |
| Swe Bench Multilingual | reasoning code | 47.3 | 来源 |
| Hmmt 2025 | math | 38.8 | 来源 |
| Simpleqa | general reasoning | 31.0 | 来源 |
| Ojbench | reasoning | 27.1 | 来源 |
| Terminal Bench | reasoning code | 25.0 | 来源 |
| Hle | reasoning math | 4.7 | 来源 |
Pricing
API 价格对比
暂无 API 价格。