Benchmark

CLUEWSC

language reasoning text 多语言

CLUEWSC2020 is the Chinese version of the Winograd Schema Challenge, part of the CLUE benchmark. It focuses on pronoun disambiguation and coreference resolution, requiring models to determine which noun a pronoun refers to in a sentence. The dataset contains 1,244 training samples and 304 development samples extracted from contemporary Chinese literature.

语言EN
满分1
参评模型2

模型排名

名次 模型 机构 分数 来源
1 Kimi-k1.5 Moonshot AI 91.4 来源 ↗
2 DeepSeek-V3 DeepSeek 90.9 来源 ↗