The gpt-oss-20b model (technically 20.9B parameters) achieves near-parity with OpenAI o4-mini on core reasoning benchmarks, while running efficiently on a single 80 GB GPU. The gpt-oss-20b model delivers similar results to OpenAI o3‑mini on common benchmarks and can run on edge devices with just 16 GB of memory, making it ideal for on-device use cases, local inference, or rapid iteration without costly infrastructure. Both models also perform strongly on tool use, few-shot function calling, CoT reasoning (as seen in results on the Tau-Bench agentic evaluation suite) and HealthBench (even outperforming proprietary models like OpenAI o1 and GPT‑4o). Note: While referred to as '20b' for simplicity, it technically has 20.9B parameters.
Benchmarks
评测成绩
| 评测基准 | 类别 | 分数 | 来源 |
|---|---|---|---|
| MMLU benchmark | general reasoning language math | 85.3 | 来源 |
| Codeforces Competition code | math reasoning | 83.9 | 来源 |
| Codeforces Competition code | math reasoning | 74.3 | 来源 |
| GPQA | reasoning general | 71.5 | 来源 |
| TAU-bench Retail benchmark | reasoning communication | 54.8 | 来源 |
| HealthBench - Realistic health conversations | healthcare | 42.5 | 来源 |
| Humanity's Last Exam | general | 17.3 | 来源 |
| Humanity's Last Exam | general | 10.9 | 来源 |
| HealthBench Hard - Challenging health conversations | healthcare | 10.8 | 来源 |
Pricing
API 价格对比
| 服务商 | 输入价 | 输出价 | 上下文 | 吞吐(tok/s) | 延迟(s) | 函数调用 | 代码执行 | 联网搜索 |
|---|---|---|---|---|---|---|---|---|
| Kenari | $0.00 | $0.00 | 131K | — | — | ✓ | ✗ | ✗ |
| LLM Gateway | $0.04 | $0.15 | 131K | — | — | ✓ | ✗ | ✗ |
| Helicone | $0.05 | $0.20 | 131K | — | — | ✓ | ✗ | ✗ |
| Neon | $0.05 | $0.20 | 131K | — | — | ✓ | ✗ | ✗ |
| OVHcloud AI Endpoints | $0.05 | $0.18 | 131K | — | — | ✓ | ✗ | ✗ |
| FrogBot | $0.07 | $0.20 | 131K | — | — | ✓ | ✗ | ✗ |
| Regolo AI | $0.40 | $1.80 | 128K | — | — | ✓ | ✗ | ✗ |
价格单位:美元/百万 token,数据来自社区整理,仅供参考。