A research preview model focused on mathematical and logical reasoning capabilities, demonstrating improved performance on tasks requiring step-by-step reasoning, mathematical problem-solving, and code generation. The model shows enhanced capabilities in formal reasoning while maintaining strong general capabilities.
发布日期2024年12月17日
参数规模—
上下文长度200K
许可证Proprietary
知识截止—
Benchmarks
评测成绩
| 评测基准 | 类别 | 分数 | 来源 |
|---|---|---|---|
| GSM8k | math reasoning | 97.1 | 来源 |
| MATH | math reasoning | 96.4 | 来源 |
| GPQA Physics | reasoning physics | 92.8 | 来源 |
| MMLU | general reasoning language math | 91.8 | 来源 |
| MGSM | math reasoning | 89.3 | 来源 |
| HumanEval | reasoning code | 88.1 | 来源 |
| MMMLU | language reasoning math general | 87.7 | 来源 |
| GPQA | reasoning general | 78.0 | 来源 |
| MMMU | multimodal reasoning general | 77.6 | 来源 |
| AIME 2024 | math reasoning | 74.3 | 来源 |
| MathVista | math vision multimodal | 71.8 | 来源 |
| TAU-bench Retail | reasoning communication | 70.8 | 来源 |
| GPQA Biology | reasoning general | 69.2 | 来源 |
| LiveBench | math reasoning general | 67.0 | 来源 |
| GPQA Chemistry | reasoning chemistry | 64.7 | 来源 |
| TAU-bench Airline | reasoning communication | 50.0 | 来源 |
| SimpleQA | general reasoning | 47.0 | 来源 |
| SWE-Bench Verified | reasoning frontend_development code | 41.0 | 来源 |
| FrontierMath | math reasoning | 5.5 | 来源 |
Pricing
API 价格对比
| 服务商 | 输入价 | 输出价 | 上下文 | 吞吐(tok/s) | 延迟(s) | 函数调用 | 代码执行 | 联网搜索 |
|---|---|---|---|---|---|---|---|---|
| Azure | $15.00 | $60.00 | 200K | — | — | ✓ | ✗ | ✗ |
| Azure Cognitive Services | $15.00 | $60.00 | 200K | — | — | ✓ | ✗ | ✗ |
| Helicone | $15.00 | $60.00 | 200K | — | — | ✗ | ✗ | ✗ |
| LLM Gateway | $15.00 | $60.00 | 200K | — | — | ✓ | ✗ | ✗ |
| OpenAI | $15.00 | $60.00 | 200K | — | — | ✓ | ✗ | ✗ |
价格单位:美元/百万 token,数据来自社区整理,仅供参考。