Benchmark
Multi-Challenge
communication
reasoning
text
MultiChallenge is a realistic multi-turn conversation evaluation benchmark that challenges frontier LLMs across four key categories: instruction retention (maintaining instructions throughout conversations), inference memory (recalling and connecting details from previous turns), reliable versioned editing (adapting to evolving instructions during collaborative editing), and self-coherence (avoiding contradictions in responses). The benchmark evaluates models on sustained, contextually complex dialogues across diverse topics including travel planning, technical documentation, and professional communication.
语言EN
满分1
参评模型7
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Kimi K2 Instruct | Moonshot AI | 54.1 | 来源 ↗ |
| 2 | Kimi K2-Instruct-0905 | Moonshot AI | 54.1 | 来源 ↗ |
| 3 | GPT-4.5 | OpenAI | 43.8 | 来源 ↗ |
| 4 | o3-mini | OpenAI | 39.9 | 来源 ↗ |
| 5 | GPT-4.1 | OpenAI | 38.3 | 来源 ↗ |
| 6 | GPT-4.1 mini | OpenAI | 35.8 | 来源 ↗ |
| 7 | GPT-4.1 nano | OpenAI | 15.0 | 来源 ↗ |