Benchmark

IF

general text

Instruction-Following Evaluation (IFEval) benchmark for large language models, focusing on verifiable instructions with 25 types of instructions and around 500 prompts containing one or more verifiable constraints

语言EN
满分1
参评模型1

模型排名

名次 模型 机构 分数 来源
1 Mistral Small 3.2 24B Instruct Mistral AI 84.8 来源 ↗