Benchmark

Nexus

general text

NexusRaven benchmark for evaluating function calling capabilities of large language models in zero-shot scenarios across cybersecurity tools and API interactions

语言EN
满分1
参评模型4

模型排名

名次 模型 机构 分数 来源
1 Llama 3.1 405B Instruct Meta 58.7 来源 ↗
2 Llama 3.1 70B Instruct Meta 56.7 来源 ↗
3 Llama 3.1 8B Instruct Meta 38.5 来源 ↗
4 Llama 3.2 3B Instruct Meta 34.3 来源 ↗