Qwen3-Next-80B-A3B-Thinking is the thinking variant of the Qwen3-Next series, featuring the same groundbreaking architecture as the instruct model. Leveraging GSPO, it addresses stability and efficiency challenges of hybrid attention + high-sparsity MoE in RL training. It uses Hybrid Attention combining Gated DeltaNet and Gated Attention for efficient ultra-long context modeling, High-Sparsity MoE with 512 experts (10 activated + 1 shared), and Multi-Token Prediction. With 80B total parameters and only 3B activated, it demonstrates outstanding performance on complex reasoning tasks — outperforming Qwen3-30B-A3B-Thinking-2507, Qwen3-32B-Thinking, and even the proprietary Gemini-2.5-Flash-Thinking across multiple benchmarks. Architecture: 48 layers, 15T training tokens, hybrid layout of 12*(3*(Gated DeltaNet->MoE)->(Gated Attention->MoE)). Supports only thinking mode with automatic <think> tag inclusion, may generate longer thinking content.
Benchmarks
评测成绩
| 评测基准 | 类别 | 分数 | 来源 |
|---|---|---|---|
| MMLU-Redux | language reasoning math general | 92.5 | 来源 |
| IFEval | general | 88.9 | 来源 |
| AIME 2025 | math reasoning | 87.8 | 来源 |
| WritingBench | writing creativity communication | 84.6 | 来源 |
| MMLU-Pro | language reasoning math general | 82.7 | 来源 |
| INCLUDE | general | 78.9 | 来源 |
| MMLU-ProX | language reasoning math general | 78.7 | 来源 |
| MultiIF | reasoning communication language | 77.8 | 来源 |
| GPQA | reasoning general | 77.2 | 来源 |
| LiveBench 241125 | math reasoning general | 76.6 | 来源 |
| HMMT25 | math | 73.9 | 来源 |
| BFCL-v3 | general reasoning | 72.0 | 来源 |
| TAU1-Retail | reasoning communication | 69.6 | 来源 |
| LiveCodeBench v6 | reasoning general | 68.7 | 来源 |
| TAU2-Retail | communication reasoning | 67.8 | 来源 |
| Arena-Hard v2 | general reasoning creativity | 62.3 | 来源 |
| SuperGPQA | reasoning general math legal healthcare finance chemistry economics physics | 60.8 | 来源 |
| TAU2-Airline | reasoning communication | 60.5 | 来源 |
| PolyMATH | math reasoning spatial_reasoning multimodal vision | 56.3 | 来源 |
| TAU1-Airline | reasoning communication | 49.0 | 来源 |
| TAU2-Telecom | communication reasoning | 43.9 | 来源 |
| OJBench | reasoning | 29.7 | 来源 |
| CFEval | code | 20.7 | 来源 |
Pricing
API 价格对比
| 服务商 | 输入价 | 输出价 | 上下文 | 吞吐(tok/s) | 延迟(s) | 函数调用 | 代码执行 | 联网搜索 |
|---|---|---|---|---|---|---|---|---|
| Alibaba (China) | $0.14 | $1.43 | 131K | — | — | ✓ | ✗ | ✗ |
| LLM Gateway | $0.15 | $1.20 | 131K | — | — | ✓ | ✗ | ✗ |
| Cortecs | $0.16 | $1.31 | 128K | — | — | ✓ | ✗ | ✗ |
| Alibaba | $0.50 | $6.00 | 131K | — | — | ✓ | ✗ | ✗ |
价格单位:美元/百万 token,数据来自社区整理,仅供参考。