Benchmark

Multi-Challenge

communication reasoning text

MultiChallenge is a realistic multi-turn conversation evaluation benchmark that challenges frontier LLMs across four key categories: instruction retention (maintaining instructions throughout conversations), inference memory (recalling and connecting details from previous turns), reliable versioned editing (adapting to evolving instructions during collaborative editing), and self-coherence (avoiding contradictions in responses). The benchmark evaluates models on sustained, contextually complex dialogues across diverse topics including travel planning, technical documentation, and professional communication.

语言EN
满分1
参评模型7

模型排名

名次 模型 机构 分数 来源
1 Kimi K2 Instruct Moonshot AI 54.1 来源 ↗
2 Kimi K2-Instruct-0905 Moonshot AI 54.1 来源 ↗
3 GPT-4.5 OpenAI 43.8 来源 ↗
4 o3-mini OpenAI 39.9 来源 ↗
5 GPT-4.1 OpenAI 38.3 来源 ↗
6 GPT-4.1 mini OpenAI 35.8 来源 ↗
7 GPT-4.1 nano OpenAI 15.0 来源 ↗