Phi-3.5-MoE-instruct is a mixture-of-experts model with ~42B total parameters (6.6B active) and a 128K context window. It excels at reasoning, math, coding, and multilingual tasks, outperforming larger dense models in many benchmarks. It underwent a thorough safety post-training process (SFT + DPO) and is licensed under MIT. This model is ideal for scenarios where efficiency and high performance are both required, particularly in multi-lingual or reasoning-intensive tasks.
发布日期2024年8月23日
参数规模60B
上下文长度128K
许可证MIT
知识截止—
Benchmarks
评测成绩
| 评测基准 | 类别 | 分数 | 来源 |
|---|---|---|---|
| ARC-C | reasoning general | 91.0 | 来源 |
| OpenBookQA | reasoning general | 89.6 | 来源 |
| GSM8k | math reasoning | 88.7 | 来源 |
| PIQA | reasoning physics general | 88.6 | 来源 |
| RULER | long_context reasoning | 87.1 | 来源 |
| RepoQA | long_context reasoning code | 85.0 | 来源 |
| BoolQ | language reasoning | 84.6 | 来源 |
| HellaSwag | reasoning | 83.8 | 来源 |
| MEGA XStoryCloze | reasoning language | 82.8 | 来源 |
| Winogrande | reasoning language | 81.3 | 来源 |
| MBPP | reasoning general | 80.8 | 来源 |
| BIG-Bench Hard | reasoning math language | 79.1 | 来源 |
| MMLU | general reasoning language math | 78.9 | 来源 |
| Social IQa | reasoning psychology | 78.0 | 来源 |
| TruthfulQA | general reasoning legal healthcare finance | 77.5 | 来源 |
| MEGA XCOPA | reasoning language | 76.6 | 来源 |
| HumanEval | reasoning code | 70.7 | 来源 |
| MMMLU | language reasoning math general | 69.9 | 来源 |
| MEGA TyDi QA | language reasoning | 67.1 | 来源 |
| MEGA MLQA | language reasoning | 65.3 | 来源 |
| MEGA UDPOS | language | 60.4 | 来源 |
| MATH | math reasoning | 59.5 | 来源 |
| MGSM | math reasoning | 58.7 | 来源 |
| MMLU-Pro | language reasoning math general | 45.3 | 来源 |
| Qasper | reasoning long_context | 40.0 | 来源 |
| Arena Hard | general reasoning creativity | 37.9 | 来源 |
| GPQA | reasoning general | 36.8 | 来源 |
| GovReport | summarization long_context | 26.4 | 来源 |
| SQuALITY | summarization long_context language | 24.1 | 来源 |
| QMSum | summarization long_context | 19.9 | 来源 |
| SummScreenFD | summarization long_context | 16.9 | 来源 |
Pricing
API 价格对比
| 服务商 | 输入价 | 输出价 | 上下文 | 吞吐(tok/s) | 延迟(s) | 函数调用 | 代码执行 | 联网搜索 |
|---|---|---|---|---|---|---|---|---|
| Azure | $0.16 | $0.64 | 128K | — | — | ✗ | ✗ | ✗ |
| Azure Cognitive Services | $0.16 | $0.64 | 128K | — | — | ✗ | ✗ | ✗ |
价格单位:美元/百万 token,数据来自社区整理,仅供参考。