Benchmark

API-Bank

reasoning text

A comprehensive benchmark for tool-augmented LLMs that evaluates API planning, retrieval, and calling capabilities. Contains 314 tool-use dialogues with 753 API calls across 73 API tools, designed to assess how effectively LLMs can utilize external tools and overcome obstacles in tool leveraging.

语言EN
满分1
参评模型3

模型排名

名次 模型 机构 分数 来源
1 Llama 3.1 405B Instruct Meta 92.0 来源 ↗
2 Llama 3.1 70B Instruct Meta 90.0 来源 ↗
3 Llama 3.1 8B Instruct Meta 82.6 来源 ↗