Benchmark

BFCL-v3

general reasoning text

Berkeley Function Calling Leaderboard v3 (BFCL-v3) is an advanced benchmark that evaluates large language models' function calling capabilities through multi-turn and multi-step interactions. It introduces extended conversational exchanges where models must retain contextual information across turns and execute multiple internal function calls for complex user requests. The benchmark includes 1000 test cases across domains like vehicle control, trading bots, travel booking, and file system management, using state-based evaluation to verify both system state changes and execution path correctness.

语言EN
满分1
参评模型6

模型排名

名次 模型 机构 分数 来源
1 GLM-4.5 Zhipu AI 77.8 来源 ↗
2 GLM-4.5-Air Zhipu AI 76.4 来源 ↗
3 Qwen3-Next-80B-A3B-Thinking Alibaba Cloud / Qwen Team 72.0 来源 ↗
4 Qwen3-235B-A22B-Thinking-2507 Alibaba Cloud / Qwen Team 71.9 来源 ↗
5 Qwen3-235B-A22B-Instruct-2507 Alibaba Cloud / Qwen Team 70.9 来源 ↗
6 Qwen3-Next-80B-A3B-Instruct Alibaba Cloud / Qwen Team 70.3 来源 ↗