Grok 4, announced by xAI in summer 2025, represents a major leap in AI capabilities, described as 'the smartest AI in the world.' Built on version 6 of xAI's foundation model, it uses 100x more training compute than Grok 2 and 10x more reinforcement learning compute than Grok 3. The model achieves PhD-level performance across all academic disciplines simultaneously, scoring perfect on standardized tests like the SAT and near-perfect on graduate exams like the GRE. Unlike Grok 3, tool usage is built into the training process rather than relying on generalization. Trained using 200,000 GPUs, Grok 4 excels at complex reasoning, mathematical problem-solving, and coding tasks, though it has acknowledged weaknesses in multimodal capabilities that are being addressed in the next version.
Benchmarks
评测成绩
| 评测基准 | 类别 | 分数 | 来源 |
|---|---|---|---|
| AIME 2025 | math reasoning | 91.7 | 来源 |
| HMMT25 | math | 90.0 | 来源 |
| GPQA | reasoning general | 87.5 | 来源 |
| LiveCodeBench | reasoning general code | 79.0 | 来源 |
| Humanity's Last Exam | general | 40.0 | 来源 |
| USAMO25 | math reasoning | 37.5 | 来源 |
| ARC-AGI v2 | reasoning vision spatial_reasoning | 15.9 | 来源 |
Pricing
API 价格对比
| 服务商 | 输入价 | 输出价 | 上下文 | 吞吐(tok/s) | 延迟(s) | 函数调用 | 代码执行 | 联网搜索 |
|---|---|---|---|---|---|---|---|---|
| Helicone | $3.00 | $15.00 | 256K | — | — | ✓ | ✗ | ✗ |
| LLM Gateway | $3.00 | $15.00 | 256K | — | — | ✓ | ✗ | ✗ |
价格单位:美元/百万 token,数据来自社区整理,仅供参考。