Leaderboard

综合榜单

名次 模型 机构 分数 来源
1 DeepSeek-V3 DeepSeek 91.6 来源 ↗
2 Claude 3.5 Sonnet Anthropic 87.1 来源 ↗
3 Claude 3.5 Sonnet Anthropic 87.1 来源 ↗
4 GPT-4 Turbo OpenAI 86.0 来源 ↗
5 Nova Pro Amazon 85.4 来源 ↗
6 Llama 3.1 405B Instruct Meta 84.8 来源 ↗
7 GPT-4o OpenAI 83.4 来源 ↗
8 Claude 3 Opus Anthropic 83.1 来源 ↗
9 Claude 3.5 Haiku Anthropic 83.1 来源 ↗
10 GPT-4 OpenAI 80.9 来源 ↗
11 Nova Lite Amazon 80.2 来源 ↗
12 GPT-4o mini OpenAI 79.7 来源 ↗
13 Llama 3.1 70B Instruct Meta 79.6 来源 ↗
14 Nova Micro Amazon 79.3 来源 ↗
15 Claude 3 Sonnet Anthropic 78.9 来源 ↗
16 Claude 3 Haiku Anthropic 78.4 来源 ↗
17 Phi 4 Microsoft 75.5 来源 ↗
18 Gemini 1.5 Pro Google 74.9 来源 ↗
19 GPT-3.5 Turbo OpenAI 70.2 来源 ↗
20 Gemma 3n E4B Instructed LiteRT Preview Google 60.8 来源 ↗
21 Gemma 3n E4B Google 60.8 来源 ↗
22 Llama 3.1 8B Instruct Meta 59.5 来源 ↗
23 Granite 3.3 8B Instruct IBM 59.4 来源 ↗
24 Gemma 3n E2B Instructed LiteRT (Preview) Google 53.9 来源 ↗
25 Gemma 3n E2B Google 53.9 来源 ↗
26 IBM Granite 4.0 Tiny Preview IBM 46.2 来源 ↗
27 Granite 3.3 8B Base IBM 36.1 来源 ↗

全部榜单

agent

audio

chemistry

code

communication

creativity

economics

finance

frontend_development

general

healthcare

image-to-text

language

legal

long_context

math

multimodal

physics

psychology

reasoning

roleplay

safety

search

spatial_reasoning

speech-to-text

summarization

text-to-image

video

vision

writing