GPT-4.1 is OpenAI's latest and most advanced flagship model, significantly improving upon GPT-4 Turbo in performance across benchmarks, speed, and cost-effectiveness.
发布日期2025年4月14日
参数规模—
上下文长度1.0M
许可证Proprietary
知识截止2024年6月1日
Benchmarks
评测成绩
| 评测基准 | 类别 | 分数 | 来源 |
|---|---|---|---|
| MMLU | general reasoning language math | 90.2 | 来源 |
| CharXiv-D | reasoning vision multimodal | 87.9 | 来源 |
| IFEval | general | 87.4 | 来源 |
| MMMLU | language reasoning math general | 87.3 | 来源 |
| MMMU | multimodal reasoning general | 74.8 | 来源 |
| MathVista | math vision multimodal | 72.2 | 来源 |
| Video-MME (long, no subtitles) | vision multimodal video | 72.0 | 来源 |
| Multi-IF | reasoning communication language | 70.8 | 来源 |
| TAU-bench Retail | reasoning communication | 68.0 | 来源 |
| GPQA | reasoning general | 66.3 | 来源 |
| COLLIE | language reasoning writing | 65.8 | 来源 |
| ComplexFuncBench | long_context reasoning | 65.5 | 来源 |
| Graphwalks BFS <128k | reasoning spatial_reasoning | 61.7 | 来源 |
| Graphwalks parents <128k | reasoning spatial_reasoning | 58.0 | 来源 |
| OpenAI-MRCR: 2 needle 128k | long_context reasoning | 57.2 | 来源 |
| CharXiv-R | reasoning vision multimodal | 56.7 | 来源 |
| SWE-Bench Verified | reasoning frontend_development code | 54.6 | 来源 |
| Aider-Polyglot Edit | general code | 52.9 | 来源 |
| Aider-Polyglot | general code | 51.6 | 来源 |
| TAU-bench Airline | reasoning communication | 49.4 | 来源 |
| Internal API instruction following (hard) | general | 49.1 | 来源 |
| AIME 2024 | math reasoning | 48.1 | 来源 |
| AIME 2025 | math reasoning | 46.4 | 来源 |
| OpenAI-MRCR: 2 needle 1M | long_context reasoning | 46.3 | 来源 |
| MultiChallenge (o3-mini grader) | reasoning language | 46.2 | 来源 |
| MultiChallenge | communication reasoning | 38.3 | 来源 |
| HMMT 2025 | math | 28.9 | 来源 |
| Graphwalks parents >128k | reasoning spatial_reasoning long_context | 25.0 | 来源 |
| Graphwalks BFS >128k | reasoning spatial_reasoning long_context | 19.0 | 来源 |
| Humanity's Last Exam | general | 5.4 | 来源 |
Pricing
API 价格对比
| 服务商 | 输入价 | 输出价 | 上下文 | 吞吐(tok/s) | 延迟(s) | 函数调用 | 代码执行 | 联网搜索 |
|---|---|---|---|---|---|---|---|---|
| 302.AI | $2.00 | $8.00 | 1.0M | — | — | ✓ | ✗ | ✗ |
| Abacus | $2.00 | $8.00 | 1.0M | — | — | ✓ | ✗ | ✗ |
| Azure | $2.00 | $8.00 | 1.0M | — | — | ✓ | ✗ | ✗ |
| Azure Cognitive Services | $2.00 | $8.00 | 1.0M | — | — | ✓ | ✗ | ✗ |
| GitHub Copilot | $2.00 | $8.00 | 128K | — | — | ✓ | ✗ | ✗ |
| Helicone | $2.00 | $8.00 | 1.0M | — | — | ✓ | ✗ | ✗ |
| LLM Gateway | $2.00 | $8.00 | 1.0M | — | — | ✓ | ✗ | ✗ |
| OpenAI | $2.00 | $8.00 | 1.0M | — | — | ✓ | ✗ | ✗ |
| Pioneer | $2.00 | $8.00 | 1.0M | — | — | ✓ | ✗ | ✗ |
| SAP AI Core | $2.00 | $8.00 | 1.0M | — | — | ✓ | ✗ | ✗ |
| Cortecs | $2.35 | $9.42 | 1.0M | — | — | ✓ | ✗ | ✗ |
价格单位:美元/百万 token,数据来自社区整理,仅供参考。