Benchmark
SciCode
reasoning
math
physics
chemistry
code
text
SciCode is a research coding benchmark curated by scientists that challenges language models to code solutions for scientific problems. It contains 338 subproblems decomposed from 80 challenging main problems across 16 natural science sub-fields including mathematics, physics, chemistry, biology, and materials science. Problems require knowledge recall, reasoning, and code synthesis skills.
语言EN
满分1
参评模型18
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Fugu | Sakana AI | 60.1 | 来源 ↗ |
| 2 | Fugu Ultra | Sakana AI | 58.7 | 来源 ↗ |
| 3 | Qwen3.7 Max | Alibaba Cloud / Qwen Team | 53.5 | 来源 ↗ |
| 4 | GLM-4.5 | Zhipu AI | 41.7 | 来源 ↗ |
| 5 | Step 3.5 Flash | StepFun | 40.4 | 来源 ↗ |
| 6 | Step 3.7 Flash | StepFun | 40.0 | 来源 ↗ |
| 7 | Step 3.5 Flash 2603 | StepFun | 38.5 | 来源 ↗ |
| 8 | Qwen3 Max | Alibaba Cloud / Qwen Team | 38.3 | 来源 ↗ |
| 9 | Mistral Small 4 | Mistral AI | 38.0 | 来源 ↗ |
| 10 | GLM-4.5-Air | Zhipu AI | 37.3 | 来源 ↗ |
| 11 | Mistral Large 3 | Mistral AI | 36.2 | 来源 ↗ |
| 12 | GPT-4o (2024-11-20) | OpenAI | 33.3 | 来源 ↗ |
| 13 | Mistral Medium 3 | Mistral AI | 33.1 | 来源 ↗ |
| 14 | Devstral 2 | Mistral AI | 33.1 | 来源 ↗ |
| 15 | Mistral Large 2.1 | Mistral AI | 29.2 | 来源 ↗ |
| 16 | Qwen3-Coder 30B-A3B Instruct | Alibaba Cloud / Qwen Team | 27.8 | 来源 ↗ |
| 17 | Sonar | Perplexity | 22.9 | 来源 ↗ |
| 18 | Sonar Pro | Perplexity | 22.6 | 来源 ↗ |