Benchmark

Tau2 Telecom

communication reasoning text

τ²-Bench telecom domain evaluates conversational agents in a dual-control environment modeled as a Dec-POMDP, where both agent and user use tools in shared telecommunications troubleshooting scenarios that test coordination and communication capabilities.

语言EN
满分1
参评模型9

模型排名

名次 模型 机构 分数 来源
1 GPT-5 OpenAI 96.7 来源 ↗
2 Claude Haiku 4.5 Anthropic 83.0 来源 ↗
3 Kimi K2 Instruct Moonshot AI 65.8 来源 ↗
4 Kimi K2-Instruct-0905 Moonshot AI 65.8 来源 ↗
5 o3 OpenAI 58.2 来源 ↗
6 Qwen3-235B-A22B-Thinking-2507 Alibaba Cloud / Qwen Team 45.6 来源 ↗
7 Qwen3-Next-80B-A3B-Thinking Alibaba Cloud / Qwen Team 43.9 来源 ↗
8 GPT-4o OpenAI 23.5 来源 ↗
9 Qwen3-Next-80B-A3B-Instruct Alibaba Cloud / Qwen Team 13.2 来源 ↗