Benchmark

Scale MultiChallenge

reasoning communication general text

MultiChallenge is a realistic multi-turn conversation evaluation benchmark developed by Scale AI that evaluates large language models on four challenging conversation categories: instruction retention, inference memory of user information, reliable versioned editing, and self-coherence. Each challenge requires accurate instruction-following, context allocation, and in-context reasoning. Despite achieving near-perfect scores on existing multi-turn evaluation benchmarks, all frontier models have less than 50% accuracy on MultiChallenge.

语言EN
满分1
参评模型4

模型排名

名次 模型 机构 分数 来源
1 GPT-5 OpenAI 69.6 来源 ↗
2 o3 OpenAI 60.4 来源 ↗
3 o4-mini OpenAI 43.0 来源 ↗
4 GPT-4o OpenAI 40.3 来源 ↗