Qwen3-235B-A22B-Thinking-2507 is a state-of-the-art thinking-enabled Mixture-of-Experts (MoE) model with 235B total parameters (22B activated). It features 94 layers, 128 experts (8 activated), and supports 262K native context length. This version delivers significantly improved reasoning performance, achieving state-of-the-art results among open-source thinking models on logical reasoning, mathematics, science, coding, and academic benchmarks. Key enhancements include markedly better general capabilities (instruction following, tool usage, text generation), enhanced 256K long-context understanding, and increased thinking depth. The model supports only thinking mode with automatic <think> tag inclusion.
Benchmarks
评测成绩
| 评测基准 | 类别 | 分数 | 来源 |
|---|---|---|---|
| MMLU-Redux | language reasoning math general | 93.8 | 来源 |
| AIME 2025 | math reasoning | 92.3 | 来源 |
| WritingBench | writing creativity communication | 88.3 | 来源 |
| IFEval | general | 87.8 | 来源 |
| Creative Writing v3 | creativity writing | 86.1 | 来源 |
| MMLU-Pro | language reasoning math general | 84.4 | 来源 |
| HMMT25 | math | 83.9 | 来源 |
| GPQA | reasoning general | 81.1 | 来源 |
| INCLUDE | general | 81.0 | 来源 |
| MMLU-ProX | language reasoning math general | 81.0 | 来源 |
| MultiIF | reasoning communication language | 80.6 | 来源 |
| Arena-Hard v2 | general reasoning creativity | 79.7 | 来源 |
| LiveBench 20241125 | math reasoning general | 78.4 | 来源 |
| LiveCodeBench v6 | reasoning general | 74.1 | 来源 |
| BFCL-v3 | general reasoning | 71.9 | 来源 |
| TAU2-Retail | communication reasoning | 71.9 | 来源 |
| TAU1-Retail | reasoning communication | 67.8 | 来源 |
| SuperGPQA | reasoning general math legal healthcare finance chemistry economics physics | 64.9 | 来源 |
| PolyMATH | math reasoning spatial_reasoning multimodal vision | 60.1 | 来源 |
| TAU2-Airline | reasoning communication | 58.0 | 来源 |
| TAU1-Airline | reasoning communication | 46.0 | 来源 |
| TAU2-Telecom | communication reasoning | 45.6 | 来源 |
| OJBench | reasoning | 32.5 | 来源 |
| CFEval | code | 21.3 | 来源 |
| HLE | general | 18.2 | 来源 |
Pricing
API 价格对比
| 服务商 | 输入价 | 输出价 | 上下文 | 吞吐(tok/s) | 延迟(s) | 函数调用 | 代码执行 | 联网搜索 |
|---|---|---|---|---|---|---|---|---|
| iFlow | $0.00 | $0.00 | 256K | — | — | ✓ | ✗ | ✗ |
| LLM Gateway | $0.20 | $0.60 | 262K | — | — | ✓ | ✗ | ✗ |
| Venice AI | $0.45 | $3.50 | 128K | — | — | ✓ | ✗ | ✗ |
价格单位:美元/百万 token,数据来自社区整理,仅供参考。