Benchmark
ACEBench
general
reasoning
text
ACEBench is a comprehensive benchmark for evaluating Large Language Models' tool usage capabilities across three primary evaluation types: Normal (basic tool usage scenarios), Special (tool usage with ambiguous or incomplete instructions), and Agent (multi-agent interactions simulating real-world dialogues). The benchmark covers 4,538 APIs across 8 major domains and 68 sub-domains including technology, finance, entertainment, society, health, culture, and environment, supporting both English and Chinese languages.
语言EN
满分1
参评模型2
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Kimi K2 Instruct | Moonshot AI | 76.5 | 来源 ↗ |
| 2 | Kimi K2-Instruct-0905 | Moonshot AI | 76.5 | 来源 ↗ |