Leaderboard

综合榜单

名次 模型 机构 分数 来源
1 o3 OpenAI 86.8 来源 ↗
2 o4-mini OpenAI 84.3 来源 ↗
3 Kimi-k1.5 Moonshot AI 74.9 来源 ↗
4 Llama 4 Maverick Meta 73.7 来源 ↗
5 GPT-4.1 mini OpenAI 73.1 来源 ↗
6 GPT-4.5 OpenAI 72.3 来源 ↗
7 GPT-4.1 OpenAI 72.2 来源 ↗
8 o1 OpenAI 71.8 来源 ↗
9 QvQ-72B-Preview Alibaba Cloud / Qwen Team 71.4 来源 ↗
10 Llama 4 Scout Meta 70.7 来源 ↗
11 Pixtral Large Mistral AI 69.4 来源 ↗
12 Grok-2 xAI 69.0 来源 ↗
13 Grok-2 mini xAI 68.1 来源 ↗
14 Gemini 1.5 Pro Google 68.1 来源 ↗
15 Qwen2.5-Omni-7B Alibaba Cloud / Qwen Team 67.9 来源 ↗
16 Claude 3.5 Sonnet Anthropic 67.7 来源 ↗
17 Mistral Small 3.2 24B Instruct Mistral AI 67.1 来源 ↗
18 Gemini 1.5 Flash Google 65.8 来源 ↗
19 GPT-4o OpenAI 63.8 来源 ↗
20 DeepSeek VL2 DeepSeek 62.8 来源 ↗
21 Phi-4-multimodal-instruct Microsoft 62.4 来源 ↗
22 GPT-4o OpenAI 61.4 来源 ↗
23 DeepSeek VL2 Small DeepSeek 60.7 来源 ↗
24 Pixtral-12B Mistral AI 58.0 来源 ↗
25 Llama 3.2 90B Instruct Meta 57.3 来源 ↗
26 GPT-4o mini OpenAI 56.7 来源 ↗
27 GPT-4.1 nano OpenAI 56.2 来源 ↗
28 Gemini 1.5 Flash 8B Google 54.7 来源 ↗
29 DeepSeek VL2 Tiny DeepSeek 53.6 来源 ↗
30 Grok-1.5V xAI 52.8 来源 ↗
31 Grok-1.5 xAI 52.8 来源 ↗
32 Llama 3.2 11B Instruct Meta 51.5 来源 ↗
33 Gemini 1.0 Pro Google 46.6 来源 ↗
34 Phi-3.5-vision-instruct Microsoft 43.9 来源 ↗
35 GPT-3.5 Turbo OpenAI 0.0 来源 ↗

全部榜单

agent

audio

chemistry

code

communication

creativity

economics

finance

frontend_development

general

healthcare

image-to-text

language

legal

long_context

math

multimodal

physics

psychology

reasoning

roleplay

safety

search

spatial_reasoning

speech-to-text

summarization

text-to-image

video

vision

writing