A research preview model focused on mathematical and logical reasoning capabilities, demonstrating improved performance on tasks requiring step-by-step reasoning, mathematical problem-solving, and code generation. The model shows enhanced capabilities in formal reasoning while maintaining strong general capabilities.
发布日期2024年9月12日
参数规模—
上下文长度—
许可证Proprietary
知识截止—
Benchmarks
评测成绩
| 评测基准 | 类别 | 分数 | 来源 |
|---|---|---|---|
| MGSM | math reasoning | 90.8 | 来源 |
| MMLU | general reasoning language math | 90.8 | 来源 |
| MATH | math reasoning | 85.5 | 来源 |
| GPQA | reasoning general | 73.3 | 来源 |
| LiveBench | math reasoning general | 52.3 | 来源 |
| SimpleQA | general reasoning | 42.4 | 来源 |
| AIME 2024 | math reasoning | 42.0 | 来源 |
| SWE-Bench Verified | reasoning frontend_development code | 41.3 | 来源 |
Pricing
API 价格对比
暂无 API 价格。