Benchmark
MultiChallenge (o3-mini grader)
reasoning
language
text
A realistic multi-turn conversation evaluation benchmark that challenges frontier LLMs across four key areas: instruction retention, inference memory, reliable versioned editing, and self-coherence. Despite near-perfect scores on existing benchmarks, frontier models achieve less than 50% accuracy on MultiChallenge.
语言EN
满分1
参评模型7