Benchmark
DS-Arena-Code
reasoning
text
Data Science Arena Code benchmark for evaluating LLMs on realistic data science code generation tasks. Tests capabilities in complex data processing, analysis, and programming across popular Python libraries used in data science workflows.
语言EN
满分1
参评模型1
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | DeepSeek-V2.5 | DeepSeek | 63.1 | 来源 ↗ |