Benchmark

SWE-bench Multilingual

reasoning code text 多语言

A multilingual benchmark for issue resolving in software engineering that covers Java, TypeScript, JavaScript, Go, Rust, C, and C++. Contains 1,632 high-quality instances carefully annotated from 2,456 candidates by 68 expert annotators, designed to evaluate Large Language Models across diverse software ecosystems beyond Python.

语言EN
满分1
参评模型14

模型排名

名次 模型 机构 分数 来源
1 Ornith 1.0 397B DeepReinforce 78.9 来源 ↗
2 Qwen3.7 Max Alibaba Cloud / Qwen Team 78.3 来源 ↗
3 Claude Sonnet 5 Anthropic 78.3 来源 ↗
4 Grok 4.5 xAI 78.0 来源 ↗
5 LongCat-2.0 Meituan 77.3 来源 ↗
6 Ornith 1.0 35B DeepReinforce 69.3 来源 ↗
7 Nemotron 3 Ultra 550B A55B NVIDIA 67.7 来源 ↗
8 Laguna XS 2.1 Poolside 63.1 来源 ↗
9 DeepSeek-V3.2-Exp DeepSeek 57.9 来源 ↗
10 DeepSeek-V3.1 DeepSeek 54.5 来源 ↗
11 Ornith 1.0 9B DeepReinforce 52.0 来源 ↗
12 Kimi K2 Instruct Moonshot AI 47.3 来源 ↗
13 Kimi K2-Instruct-0905 Moonshot AI 47.3 来源 ↗
14 DeepSeek-R1-0528 DeepSeek 30.5 来源 ↗