Benchmark

Tau2 Retail

communication reasoning text

τ²-bench retail domain evaluates conversational AI agents in customer service scenarios within a dual-control environment where both agent and user can interact with tools. Tests tool-agent-user interaction, rule adherence, and task consistency in retail customer support contexts.

语言EN
满分1
参评模型10

模型排名

名次 模型 机构 分数 来源
1 Claude Haiku 4.5 Anthropic 83.2 来源 ↗
2 GPT-5 OpenAI 81.1 来源 ↗
3 o3 OpenAI 80.2 来源 ↗
4 Qwen3-235B-A22B-Thinking-2507 Alibaba Cloud / Qwen Team 71.9 来源 ↗
5 Qwen3-235B-A22B-Instruct-2507 Alibaba Cloud / Qwen Team 71.3 来源 ↗
6 Kimi K2 Instruct Moonshot AI 70.6 来源 ↗
7 Kimi K2-Instruct-0905 Moonshot AI 70.6 来源 ↗
8 Qwen3-Next-80B-A3B-Thinking Alibaba Cloud / Qwen Team 67.8 来源 ↗
9 GPT-4o OpenAI 63.4 来源 ↗
10 Qwen3-Next-80B-A3B-Instruct Alibaba Cloud / Qwen Team 57.3 来源 ↗