Benchmark

Tau-bench

general reasoning text

τ-bench: A benchmark for tool-agent-user interaction in real-world domains. Tests language agents' ability to interact with users and follow domain-specific rules through dynamic conversations using API tools and policy guidelines across retail and airline domains. Evaluates consistency and reliability of agent behavior over multiple trials.

语言EN
满分1
参评模型1

模型排名

名次 模型 机构 分数 来源
1 o3 OpenAI 63.0 来源 ↗