Benchmark
BigCodeBench
general
reasoning
text
A benchmark that challenges LLMs to invoke multiple function calls as tools from 139 libraries and 7 domains for 1,140 fine-grained programming tasks. Evaluates code generation with diverse function calls and complex instructions, featuring two variants: Complete (code completion based on comprehensive docstrings) and Instruct (generating code from natural language instructions).
语言EN
满分1
参评模型2
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Gemini Diffusion | 45.4 | 来源 ↗ | |
| 2 | Qwen2.5-Coder 7B Instruct | Alibaba Cloud / Qwen Team | 41.0 | 来源 ↗ |