Leaderboard

综合榜单

名次 模型 机构 分数 来源
1 Claude 3 Opus Anthropic 95.4 来源 ↗
2 GPT-4 OpenAI 95.3 来源 ↗
3 Gemini 1.5 Pro Google 93.3 来源 ↗
4 Claude 3 Sonnet Anthropic 89.0 来源 ↗
5 Command R+ Cohere 88.6 来源 ↗
6 Qwen2 72B Instruct Alibaba Cloud / Qwen Team 87.6 来源 ↗
7 Gemini 1.5 Flash Google 86.5 来源 ↗
8 Gemma 2 27B Google 86.4 来源 ↗
9 Claude 3 Haiku Anthropic 85.9 来源 ↗
10 Llama 3.1 Nemotron 70B Instruct NVIDIA 85.6 来源 ↗
11 Qwen2.5 32B Instruct Alibaba Cloud / Qwen Team 85.2 来源 ↗
12 Phi-3.5-MoE-instruct Microsoft 83.8 来源 ↗
13 Mistral NeMo Instruct Mistral AI 83.5 来源 ↗
14 Qwen2.5-Coder 32B Instruct Alibaba Cloud / Qwen Team 83.0 来源 ↗
15 Gemma 2 9B Google 81.9 来源 ↗
16 Granite 3.3 8B Base IBM 80.1 来源 ↗
17 Gemma 3n E4B Instructed LiteRT Preview Google 78.6 来源 ↗
18 Gemma 3n E4B Google 78.6 来源 ↗
19 Qwen2.5-Coder 7B Instruct Alibaba Cloud / Qwen Team 76.8 来源 ↗
20 Gemma 3n E2B Instructed LiteRT (Preview) Google 72.2 来源 ↗
21 Gemma 3n E2B Google 72.2 来源 ↗
22 Llama 3.2 3B Instruct Meta 69.8 来源 ↗
23 Phi-3.5-mini-instruct Microsoft 69.4 来源 ↗
24 Phi 4 Mini Microsoft 69.1 来源 ↗

全部榜单

agent

audio

chemistry

code

communication

creativity

economics

finance

frontend_development

general

healthcare

image-to-text

language

legal

long_context

math

multimodal

physics

psychology

reasoning

roleplay

safety

search

spatial_reasoning

speech-to-text

summarization

text-to-image

video

vision

writing