Benchmark
Tau2 Airline
reasoning
communication
text
TAU2 airline domain benchmark for evaluating conversational agents in dual-control environments where both AI agents and users interact with tools in airline customer service scenarios. Tests agent coordination, communication, and ability to guide user actions in tasks like flight booking, modifications, cancellations, and refunds.
语言EN
满分1
参评模型10
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | o3 | OpenAI | 64.8 | 来源 ↗ |
| 2 | Claude Haiku 4.5 | Anthropic | 63.6 | 来源 ↗ |
| 3 | GPT-5 | OpenAI | 62.6 | 来源 ↗ |
| 4 | Qwen3-Next-80B-A3B-Thinking | Alibaba Cloud / Qwen Team | 60.5 | 来源 ↗ |
| 5 | Qwen3-235B-A22B-Thinking-2507 | Alibaba Cloud / Qwen Team | 58.0 | 来源 ↗ |
| 6 | Kimi K2 Instruct | Moonshot AI | 56.5 | 来源 ↗ |
| 7 | Kimi K2-Instruct-0905 | Moonshot AI | 56.5 | 来源 ↗ |
| 8 | GPT-4o | OpenAI | 45.5 | 来源 ↗ |
| 9 | Qwen3-Next-80B-A3B-Instruct | Alibaba Cloud / Qwen Team | 45.5 | 来源 ↗ |
| 10 | Qwen3-235B-A22B-Instruct-2507 | Alibaba Cloud / Qwen Team | 44.0 | 来源 ↗ |