Model

Qwen3-Next-80B-A3B-Thinking

Alibaba Cloud / Qwen Team
开源 Apache 2.0

Qwen3-Next-80B-A3B-Thinking is the thinking variant of the Qwen3-Next series, featuring the same groundbreaking architecture as the instruct model. Leveraging GSPO, it addresses stability and efficiency challenges of hybrid attention + high-sparsity MoE in RL training. It uses Hybrid Attention combining Gated DeltaNet and Gated Attention for efficient ultra-long context modeling, High-Sparsity MoE with 512 experts (10 activated + 1 shared), and Multi-Token Prediction. With 80B total parameters and only 3B activated, it demonstrates outstanding performance on complex reasoning tasks — outperforming Qwen3-30B-A3B-Thinking-2507, Qwen3-32B-Thinking, and even the proprietary Gemini-2.5-Flash-Thinking across multiple benchmarks. Architecture: 48 layers, 15T training tokens, hybrid layout of 12*(3*(Gated DeltaNet->MoE)->(Gated Attention->MoE)). Supports only thinking mode with automatic <think> tag inclusion, may generate longer thinking content.

发布日期2025年9月10日
参数规模80B
上下文长度131K
许可证Apache 2.0
知识截止

Benchmarks

评测成绩

评测基准 类别 分数 来源
MMLU-Redux language reasoning math general 92.5 来源
IFEval general 88.9 来源
AIME 2025 math reasoning 87.8 来源
WritingBench writing creativity communication 84.6 来源
MMLU-Pro language reasoning math general 82.7 来源
INCLUDE general 78.9 来源
MMLU-ProX language reasoning math general 78.7 来源
MultiIF reasoning communication language 77.8 来源
GPQA reasoning general 77.2 来源
LiveBench 241125 math reasoning general 76.6 来源
HMMT25 math 73.9 来源
BFCL-v3 general reasoning 72.0 来源
TAU1-Retail reasoning communication 69.6 来源
LiveCodeBench v6 reasoning general 68.7 来源
TAU2-Retail communication reasoning 67.8 来源
Arena-Hard v2 general reasoning creativity 62.3 来源
SuperGPQA reasoning general math legal healthcare finance chemistry economics physics 60.8 来源
TAU2-Airline reasoning communication 60.5 来源
PolyMATH math reasoning spatial_reasoning multimodal vision 56.3 来源
TAU1-Airline reasoning communication 49.0 来源
TAU2-Telecom communication reasoning 43.9 来源
OJBench reasoning 29.7 来源
CFEval code 20.7 来源

Pricing

API 价格对比

服务商 输入价 输出价 上下文 吞吐(tok/s) 延迟(s) 函数调用 代码执行 联网搜索
Alibaba (China) $0.14 $1.43 131K
LLM Gateway $0.15 $1.20 131K
Cortecs $0.16 $1.31 128K
Alibaba $0.50 $6.00 131K

价格单位:美元/百万 token,数据来自社区整理,仅供参考。