GPT-4o ('o' for 'omni') is a multimodal AI model that accepts text, audio, image, and video inputs, and generates text, audio, and image outputs. It matches GPT-4 Turbo performance on text and code, with improvements in non-English languages, vision, and audio understanding.
发布日期2024年8月6日
参数规模—
上下文长度128K
许可证Proprietary
知识截止—
Benchmarks
评测成绩
| 评测基准 | 类别 | 分数 | 来源 |
|---|---|---|---|
| AI2D | vision reasoning multimodal | 94.2 | 来源 |
| DocVQA | vision multimodal | 92.8 | 来源 |
| MMLU | general reasoning language math | 85.7 | 来源 |
| ChartQA | reasoning vision multimodal | 85.7 | 来源 |
| CharXiv-D | reasoning vision multimodal | 85.3 | 来源 |
| MMMLU | language reasoning math general | 81.4 | 来源 |
| IFEval | general | 81.0 | 来源 |
| MMLU-Pro | language reasoning math general | 74.7 | 来源 |
| EgoSchema | vision reasoning long_context | 72.2 | 来源 |
| MMMU | multimodal reasoning general | 72.2 | 来源 |
| GPQA | reasoning general | 70.1 | 来源 |
| ComplexFuncBench | long_context reasoning | 66.5 | 来源 |
| Tau2 retail | communication reasoning | 63.4 | 来源 |
| ActivityNet | vision video | 61.9 | 来源 |
| MathVista | math vision multimodal | 61.4 | 来源 |
| VideoMMMU | multimodal vision reasoning | 61.2 | 来源 |
| COLLIE | language reasoning writing | 61.0 | 来源 |
| Multi-IF | reasoning communication language | 60.9 | 来源 |
| TAU-bench Retail | reasoning communication | 60.3 | 来源 |
| MMMU-Pro | vision multimodal reasoning general | 59.9 | 来源 |
| CharXiv-R | reasoning vision multimodal | 58.8 | 来源 |
| Tau2 airline | reasoning communication | 45.5 | 来源 |
| TAU-bench Airline | reasoning communication | 42.8 | 来源 |
| Graphwalks BFS <128k | reasoning spatial_reasoning | 41.7 | 来源 |
| Scale MultiChallenge | reasoning communication general | 40.3 | 来源 |
| MultiChallenge (o3-mini grader) | reasoning language | 39.9 | 来源 |
| SimpleQA | general reasoning | 38.2 | 来源 |
| Graphwalks parents <128k | reasoning spatial_reasoning | 35.4 | 来源 |
| ERQA | vision reasoning spatial_reasoning | 35.2 | 来源 |
| SWE-Bench Verified | reasoning frontend_development code | 33.2 | 来源 |
| SWE-Lancer | reasoning code | 32.6 | 来源 |
| OpenAI-MRCR: 2 needle 128k | long_context reasoning | 31.9 | 来源 |
| Aider-Polyglot | general code | 30.7 | 来源 |
| Internal API instruction following (hard) | general | 29.2 | 来源 |
| Tau2 telecom | communication reasoning | 23.5 | 来源 |
| Aider-Polyglot Edit | general code | 18.2 | 来源 |
| AIME 2024 | math reasoning | 13.1 | 来源 |
| SWE-Lancer (IC-Diamond subset) | reasoning code | 12.4 | 来源 |
| Humanity's Last Exam | general | 5.3 | 来源 |
Pricing
API 价格对比
| 服务商 | 输入价 | 输出价 | 上下文 | 吞吐(tok/s) | 延迟(s) | 函数调用 | 代码执行 | 联网搜索 |
|---|---|---|---|---|---|---|---|---|
| OpenAI | $2.50 | $10.00 | 128K | — | — | ✓ | ✗ | ✗ |
价格单位:美元/百万 token,数据来自社区整理,仅供参考。