What's New

新模型动态

2026年7月

新发布 开源
P

Laguna S 2.1

Poolside

Agentic coding model from Poolside in the XS size class for local deployment

2026年7月21日
新发布 专有 多模态

Fast Gemini model balancing multimodal reasoning, tool use, and cost

2026年7月21日 1.0M 上下文 $0.30 起
新发布 专有 多模态
G

Fast Gemini model balancing multimodal reasoning, tool use, and cost

2026年7月21日 1.0M 上下文 $1.50 起
新发布 专有 多模态
A

Qwen3.8 Max Preview

Alibaba Cloud / Qwen Team

Preview Qwen flagship for million-token multimodal reasoning and long-horizon agentic workflows

2026年7月19日 1.0M 上下文 $0.00 起
新发布 开源 多模态
M

Kimi K3

Moonshot AI

Multimodal Kimi model with 1M context and toggleable max-effort thinking for long-horizon agent work

2026年7月16日 1.0M 上下文 $0.00 起
新发布 开源 多模态
T

Inkling

Thinking Machines Lab

Multimodal MoE reasoning model (975B total, 41B active) for text, image, and audio

2026年7月15日 1.0M 上下文 $1.25 起
新发布 专有 多模态
O

GPT-5.6 Luna

OpenAI

Cost-efficient GPT-5.6 model for fast, high-volume workloads

2026年7月9日 1.0M 上下文 $0.70 起
新发布 专有 多模态
O

GPT-5.6 Sol

OpenAI

Frontier GPT-5.6 model for complex professional work, coding, and agentic workflows

2026年7月9日 1.0M 上下文 $2.50 起
新发布 专有 多模态
O

GPT-5.6 Terra

OpenAI

Balanced GPT-5.6 model for capable, cost-efficient everyday work

2026年7月9日 1.0M 上下文 $1.50 起
新发布 专有 多模态
X

Grok 4.5

xAI

xAI's latest Grok for chat, coding, agentic tools, and lower hallucination risk

2026年7月8日 1.0M 上下文 $2.00 起
新发布 专有 多模态
O

Realtime speech-to-speech model with configurable reasoning, tool use, and robust voice-agent behavior

2026年7月6日 96K 上下文 $4.00 起
新发布 开源
T

Hy3

Tencent

Tencent Hy reasoning model for coding, instruction following, and agent tasks

2026年7月6日 262K 上下文 $0.00 起
新发布 开源
P

Laguna XS 2.1

Poolside

Agentic coding model from Poolside in the XS size class for local deployment

2026年7月2日

2026年6月

新发布 专有 多模态
A

Claude Sonnet 5

Anthropic

Everyday Claude agent model for coding, planning, browsing, and general work

2026年6月30日 1.0M 上下文 $0.00 起
新发布 专有 多模态

Video generation and editing model for fast, conversational text- and image-to-video workflows

2026年6月30日 131K 上下文 $1.50 起
新发布 专有
M

LongCat-2.0

Meituan

Meituan LongCat-2.0, a reasoning model with tool calling and a 1M-token context window

2026年6月30日 1.0M 上下文 $0.75 起
新发布 开源 多模态
D

Ornith 1.0 31B

DeepReinforce

Open coding-reasoning model for repository tasks and self-improving agents

2026年6月25日
新发布 开源 多模态
D

Ornith 1.0 35B

DeepReinforce

Large coding-reasoning model for agentic software tasks and RL search

2026年6月25日
新发布 开源 多模态
D

Ornith 1.0 397B

DeepReinforce

Large coding-reasoning model for agentic software tasks and RL search

2026年6月25日
新发布 开源 多模态
D

Ornith 1.0 9B

DeepReinforce

Open coding-reasoning model for repository tasks and self-improving agents

2026年6月25日
专有 多模态
S

Fugu

Sakana AI

Multi-agent model for routing expert agents across complex analytical tasks

2026年6月15日
专有 多模态
S

Fugu Ultra

Sakana AI

Quality-first multi-agent model for hard research, analysis, and competitions

2026年6月15日 1.0M 上下文 $5.00 起
开源
Z

GLM-5.2

Zhipu AI

Open flagship GLM for long-horizon coding agents and million-token context work

2026年6月13日 1.0M 上下文 $0.00 起
开源 多模态
M

Kimi K2.7 Code

Moonshot AI

Coding-focused Kimi model, stronger on long-horizon repo work with less overthinking

2026年6月12日 262K 上下文 $0.00 起
开源 多模态
M

Lower-latency Kimi Code variant for interactive edits and coding-agent loops

2026年6月12日 262K 上下文 $1.90 起
专有 多模态
A

Claude Fable 5

Anthropic

Claude model for creative writing, analysis, and controlled agent workflows

2026年6月9日 1.0M 上下文 $0.00 起
开源
C

Cohere coding model for practical software engineering and agentic edits

2026年6月9日 256K 上下文 $0.00 起
开源

MiMo pro model for strong multimodal reasoning and agent execution

2026年6月8日 1.0M 上下文 $1.31 起
开源

Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy

2026年6月4日 1.0M 上下文 $0.00 起
开源 多模态

Safety model for policy screening, moderation, and risk-aware routing workflows

2026年6月4日
专有 多模态
A

Qwen3.7 Plus

Alibaba Cloud / Qwen Team

Multimodal Qwen workhorse for long-context agents, visual inputs, and coding

2026年6月2日 1.0M 上下文 $0.00 起
专有
M

MAI-Code-1-Flash

Microsoft

Microsoft coding model built for fast, efficient assistance in everyday developer workflows

2026年6月2日
开源 多模态
M

MiniMax-M3

MiniMax

MiniMax multimodal model for long-context coding, perception, and agent planning

2026年6月1日 1.0M 上下文 $0.00 起

2026年5月

开源 多模态
S

Step 3.7 Flash

StepFun

Newer StepFun flash model for faster agents, coding, and multimodal prompts

2026年5月29日 256K 上下文 $0.19 起
专有 多模态
A

Claude Opus 4.8

Anthropic

Top Claude Opus tier for the hardest reasoning, coding, and long-horizon agents

2026年5月28日 1.0M 上下文 $0.00 起
专有 多模态
G

Nano Banana Pro for higher-fidelity image generation and design-heavy edits

2026年5月28日 66K 上下文 $2.00 起
专有 多模态
G

Nano Banana 2

Google

Image model for prompt-driven generation, editing, and visual design workflows

2026年5月28日 131K 上下文 $0.50 起
专有
A

Qwen3.7 Max

Alibaba Cloud / Qwen Team

Qwen frontier model tuned for agent frameworks, coding assistants, and long tasks

2026年5月21日 1.0M 上下文 $0.00 起
开源 多模态
C

Cohere's stronger command model for multilingual agents and enterprise workflows

2026年5月20日 128K 上下文 $2.50 起
专有 多模态
G

Fast Gemini model balancing multimodal reasoning, tool use, and cost

2026年5月19日 1.0M 上下文 $0.19 起
专有 多模态

Fast Gemini model balancing multimodal reasoning, tool use, and cost

2026年5月19日 1.0M 上下文 $1.50 起
专有 多模态

Low-latency Gemini model for high-volume multimodal and agent workloads

2026年5月7日 1.0M 上下文 $0.25 起
专有 多模态

Low-latency Gemini model for high-volume multimodal and agent workloads

2026年5月7日 1.0M 上下文 $0.25 起
专有 多模态

Streaming speech-to-text model for low-latency transcript deltas from live audio

2026年5月7日
专有 多模态
O

Compact GPT model for low-latency assistance and high-volume workloads

2026年5月5日

2026年4月

开源 多模态
M

Mistral Medium 3.5

Mistral AI

Balanced Mistral model for enterprise assistants, multilingual work, and tools

2026年4月29日 262K 上下文 $1.50 起
开源 多模态

Balanced Mistral model for enterprise assistants, multilingual work, and tools

2026年4月29日 262K 上下文 $1.50 起
开源
P

Laguna M.1

Poolside

Poolside's open-weight model for agentic coding and long-horizon work

2026年4月28日
开源
P

Laguna XS.2

Poolside

Agentic coding model from Poolside in the XS size class for local deployment

2026年4月28日
专有 多模态
A

Qwen3.6 Flash

Alibaba Cloud / Qwen Team

Qwen vision-language model for visual reasoning, documents, and agent tasks

2026年4月27日 1.0M 上下文 $0.00 起
开源
D

Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work

2026年4月24日 1.1M 上下文 $0.00 起
开源
D

DeepSeek V4 Pro

DeepSeek

Open MoE flagship with million-token context for coding and long agent runs

2026年4月24日 1.1M 上下文 $0.00 起
专有 多模态
O

GPT-5.5

OpenAI

Default frontier GPT for coding, computer use, research, and knowledge work

2026年4月23日 1.1M 上下文 $0.19 起
专有 多模态
O

GPT-5.5 Pro

OpenAI

Highest-accuracy GPT-5.5 tier for slower, precision-heavy reasoning and coding

2026年4月23日 922K 上下文 $30.00 起
开源 多模态
X

MiMo-V2.5

Xiaomi

Open MiMo model for multimodal coding agents and long-context automation

2026年4月22日 1.0M 上下文 $0.00 起
开源
X

MiMo-V2.5-Pro

Xiaomi

Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution

2026年4月22日 1.0M 上下文 $0.00 起
开源 多模态
A

Qwen3.6 27B

Alibaba Cloud / Qwen Team

Qwen vision-language model for visual reasoning, documents, and agent tasks

2026年4月22日 262K 上下文 $0.20 起
开源 多模态
M

Kimi K2.6

Moonshot AI

Multimodal Kimi workhorse for agent loops, coding tasks, and visual context

2026年4月21日 262K 上下文 $0.00 起
专有 多模态
O

GPT-Image-2

OpenAI

Image model for prompt-driven generation, editing, and visual design workflows

2026年4月21日 272K 上下文 $0.00 起
开源
T

Hy3 preview

Tencent

Tencent Hy reasoning model for coding, instruction following, and agent tasks

2026年4月20日 256K 上下文 $0.00 起
专有
A

Qwen3.6 Max Preview

Alibaba Cloud / Qwen Team

Flagship Qwen model for complex reasoning, coding, and agentic workflows

2026年4月20日 262K 上下文 $1.04 起
开源 多模态
A

Qwen3.6 35B-A3B

Alibaba Cloud / Qwen Team

Open multimodal Qwen MoE for local agents that need vision, audio, and code

2026年4月17日 262K 上下文 $0.25 起
专有 多模态
X

Grok 4.3

xAI

xAI's default Grok for chat, coding, agentic tools, and lower hallucination risk

2026年4月17日 1.0M 上下文 $1.25 起
专有 多模态

Fast Grok coding model tuned for agentic engineering and iterative edits

2026年4月16日 256K 上下文 $1.00 起
专有 多模态
A

Claude Opus 4.7

Anthropic

Stronger Opus tier for advanced software work and high-stakes reasoning

2026年4月16日 1.0M 上下文 $0.00 起
开源

Safety model for policy screening, moderation, and risk-aware routing workflows

2026年4月16日
专有 多模态

Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration.

2026年4月8日 1.0M 上下文 $1.25 起
开源
Z

GLM-5.1

Zhipu AI

Strong GLM coding model for agentic engineering, terminals, and repository generation

2026年4月7日 205K 上下文 $0.00 起
开源

StepFun flash model for efficient multimodal reasoning, coding, and tool use

2026年4月2日 256K 上下文 $0.10 起
专有 多模态
A

Qwen3.6 Plus

Alibaba Cloud / Qwen Team

Earlier Qwen multimodal workhorse for million-token agent and document tasks

2026年4月2日 1.0M 上下文 $0.00 起
开源 多模态

Open Gemma instruction model for efficient chat and self-hosted deployments

2026年4月2日 262K 上下文 $0.07 起
开源 多模态
G

Largest Gemma 4 instruction model for open, self-hosted chat and reasoning

2026年4月2日 262K 上下文 $0.00 起
开源 多模态
G

Open Gemma instruction model for efficient chat and self-hosted deployments

2026年4月2日
开源 多模态
G

Open Gemma instruction model for efficient chat and self-hosted deployments

2026年4月2日
专有 多模态
Z

GLM-5V-Turbo

Zhipu AI

Fast GLM vision model for screenshots, documents, and multimodal agent tasks

2026年4月1日 200K 上下文 $0.00 起

2026年3月

开源 多模态

Reranking model for improving retrieval quality in search and recommendation systems

2026年3月31日
开源

Nemotron model for efficient reasoning, coding, and specialized AI agents

2026年3月24日
开源
M

MiniMax-M2.7

MiniMax

Open MiniMax flagship for coding agents, office automation, and complex environments

2026年3月18日 205K 上下文 $0.00 起
开源

Low-latency M2.7 variant for interactive coding plans and agent loops

2026年3月18日 205K 上下文 $0.00 起
专有 多模态
X

MiMo-V2-Omni

Xiaomi

MiMo omni model for text, image, video, audio, and agents

2026年3月18日 262K 上下文 $0.14 起
专有
X

MiMo-V2-Pro

Xiaomi

Earlier MiMo Pro model for multimodal agents, reasoning, and code tasks

2026年3月18日 1.0M 上下文 $0.00 起
专有 多模态
O

GPT-5.4 mini

OpenAI

Strong small GPT for coding subagents, quick tool use, and high-volume work

2026年3月17日 272K 上下文 $0.75 起
专有 多模态
O

GPT-5.4 nano

OpenAI

Cheapest GPT-5.4 lane for simple routing, extraction, and bulk automation

2026年3月17日 1.0M 上下文 $0.20 起
专有
Z

GLM-5-Turbo

Zhipu AI

Faster GLM-5 lane for coding agents that need lower latency

2026年3月16日 200K 上下文 $0.00 起
开源 多模态
M

Mistral Small 4

Mistral AI

Fast Mistral production model for chat, extraction, and cost-sensitive agents

2026年3月16日 256K 上下文 $0.15 起
开源 多模态

Nemotron multimodal model for visual reasoning and agentic AI workflows

2026年3月16日
开源

Nemotron middle tier for collaborative agents and high-volume reasoning workloads

2026年3月11日 262K 上下文 $0.00 起
专有 多模态

Grok model for agentic tool use, reasoning, coding, and live assistance

2026年3月9日 1.0M 上下文 $1.25 起
专有 多模态

Reasoning Grok for document-heavy analysis and long-horizon tool use

2026年3月9日 1.0M 上下文 $1.25 起
专有 多模态
O

GPT-5.4

OpenAI

Agent-ready GPT for coding and computer-use workflows at a lower cost

2026年3月5日 1.1M 上下文 $1.80 起
专有 多模态
O

GPT-5.4 Pro

OpenAI

More exact GPT-5.4 tier for demanding professional reasoning and agent tasks

2026年3月5日 922K 上下文 $30.00 起
专有 多模态

Chat-tuned GPT model for conversational assistance, writing, and tool workflows

2026年3月3日 400K 上下文 $1.75 起
专有 多模态

Low-latency Gemini model for high-volume multimodal and agent workloads

2026年3月3日 1.0M 上下文 $0.25 起

2026年2月

专有 多模态
G

Nano Banana 2

Google

Image model for prompt-driven generation, editing, and visual design workflows

2026年2月26日 1.0M 上下文 $0.50 起
开源 多模态
A

Qwen3.5 122B-A10B

Alibaba Cloud / Qwen Team

Qwen vision-language model for visual reasoning, documents, and agent tasks

2026年2月23日 262K 上下文 $0.36 起
开源 多模态
A

Qwen3.5 27B

Alibaba Cloud / Qwen Team

Qwen vision-language model for visual reasoning, documents, and agent tasks

2026年2月23日 262K 上下文 $0.27 起
开源 多模态
A

Qwen3.5 35B-A3B

Alibaba Cloud / Qwen Team

Qwen vision-language model for visual reasoning, documents, and agent tasks

2026年2月23日 262K 上下文 $0.23 起
开源
A

Qwen3.5 9B

Alibaba Cloud / Qwen Team

Qwen instruction model for multilingual chat, reasoning, and tool use

2026年2月23日 262K 上下文 $0.04 起
专有 多模态

Reasoning-first Gemini preview for agentic coding and complex problem solving

2026年2月19日 1.0M 上下文 $2.00 起
专有 多模态

Advanced Gemini model for complex reasoning, coding, and multimodal analysis

2026年2月19日 1.0M 上下文 $2.00 起
开源
S

Sarvam 30B

Sarvam AI

Efficient Indian-language reasoning model for chat, coding, and multilingual work

2026年2月18日 66K 上下文 $0.03 起
专有 多模态
A

Claude Sonnet 4.6

Anthropic

Claude workhorse for coding agents, careful analysis, and production cost control

2026年2月17日 1.0M 上下文 $3.00 起
专有 多模态
A

Qwen3.5 Plus

Alibaba Cloud / Qwen Team

Qwen vision-language model for visual reasoning, documents, and agent tasks

2026年2月16日 1.0M 上下文 $0.00 起
开源 多模态
A

Qwen3.5 397B-A17B

Alibaba Cloud / Qwen Team

Large open Qwen multimodal MoE for visual agents and long technical tasks

2026年2月15日 262K 上下文 $0.35 起
开源

High-speed MiniMax model for low-latency coding and agent workflows

2026年2月13日 205K 上下文 $0.00 起
开源
M

MiniMax-M2.5

MiniMax

Prior MiniMax coding model for agent workflows, office edits, and automation

2026年2月12日 229K 上下文 $0.00 起
开源
Z

GLM-5

Zhipu AI

General GLM flagship for coding, analysis, and tool-heavy engineering workflows

2026年2月12日 205K 上下文 $0.00 起
开源 多模态

Embedding model for semantic search, retrieval, clustering, and ranking pipelines

2026年2月10日
专有 多模态
A

Claude Opus 4.6

Anthropic

High-end Claude for difficult coding, planning, and slower expert reasoning

2026年2月5日 1.0M 上下文 $5.00 起
专有 多模态
O

GPT-5.3 Codex

OpenAI

Coding-optimized GPT model for repository edits, reviews, and agentic software work

2026年2月5日 400K 上下文 $1.75 起

2026年1月

开源
S

Step 3.5 Flash

StepFun

StepFun flash lane for quick multimodal reasoning and coding assistance

2026年1月29日 256K 上下文 $0.10 起
开源
Z

GLM-4.7-Flash

Zhipu AI

Budget GLM lane for fast coding help, routing, and everyday automation

2026年1月19日 203K 上下文 $0.00 起
开源
Z

GLM-4.7-FlashX

Zhipu AI

Efficient GLM model for fast reasoning, coding, and agent workflows

2026年1月19日 200K 上下文 $0.07 起
开源 多模态
M

Kimi K2.5

Moonshot AI

Earlier Kimi frontier model for long-context agents, coding, and multimodal work

2026年1月1日 262K 上下文 $0.00 起

2025年12月

开源
M

MiniMax-M2.1

MiniMax

Earlier MiniMax agent model for practical coding and productivity tasks

2025年12月23日 1.0M 上下文 $0.00 起
开源
Z

GLM-4.7

Zhipu AI

Mature GLM model for dependable coding, reasoning, and structured agent tasks

2025年12月22日 205K 上下文 $0.00 起
专有 多模态

New Gemini flash lane bringing frontier-style multimodal reasoning to cheaper runs

2025年12月17日 1.0M 上下文 $0.07 起
开源
X

MiMo-V2-Flash

Xiaomi

MiMo flash model for fast multimodal assistance and agent workflows

2025年12月16日 262K 上下文 $0.10 起
开源

Small Nemotron 3 MoE for efficient coding, math, and long-context agents

2025年12月15日 262K 上下文 $0.00 起
专有 多模态
O

GPT-5.2

OpenAI

Reliable GPT generation for broad coding, writing, and tool-assisted product work

2025年12月11日 400K 上下文 $0.25 起
专有 多模态
O

GPT-5.2 Chat

OpenAI

Chat-tuned GPT model for conversational assistance, writing, and tool workflows

2025年12月11日 400K 上下文 $1.75 起
专有 多模态
O

GPT-5.2 Codex

OpenAI

Code-specialist GPT for repository edits, reviews, and long-running software agents

2025年12月11日 400K 上下文 $0.14 起
专有 多模态
O

GPT-5.2 Pro

OpenAI

Higher-accuracy GPT-5.2 variant for tougher reasoning and review workflows

2025年12月11日 400K 上下文 $18.90 起
开源
M

Devstral 2

Mistral AI

Mistral's coding-agent model for repository work, terminal tasks, and software fixes

2025年12月9日 262K 上下文 $0.00 起
开源 多模态
Z

GLM-4.6V

Zhipu AI

GLM vision model for visual reasoning, documents, and multimodal agents

2025年12月8日 131K 上下文 $0.15 起
开源
M

Devstral 2 (latest)

Mistral AI

Mistral coding agent model for repository tasks and software engineering workflows

2025年12月2日 262K 上下文 $0.40 起
开源
D

DeepSeek Chat

DeepSeek

DeepSeek chat model for instruction following, coding, and analysis

2025年12月1日 1.0M 上下文 $0.14 起
开源
D

DeepSeek reasoning model for multi-step analysis, math, coding, and tools

2025年12月1日 1.0M 上下文 $0.14 起

2025年11月

专有 多模态
O

GPT-Image-1.5

OpenAI

Image model for prompt-driven generation, editing, and visual design workflows

2025年11月25日 $5.00 起
专有 多模态

Flagship Claude model for deep reasoning, coding, and long-horizon agents

2025年11月24日 200K 上下文 $5.00 起
专有 多模态
G

Nano Banana Pro for higher-fidelity image generation and design-heavy edits

2025年11月20日 1.0M 上下文 $2.00 起
专有 多模态

Preview Gemini flagship for complex reasoning, coding, and rich multimodal prompts

2025年11月18日 1.0M 上下文 $0.57 起
专有 多模态
O

GPT-5.1

OpenAI

Sharper GPT-5 generation for coding, product work, and tool-assisted tasks

2025年11月13日 400K 上下文 $1.07 起
专有 多模态
O

GPT-5.1 Chat

OpenAI

Chat-tuned GPT-5.1 for polished assistants, writing, and product conversations

2025年11月13日 400K 上下文 $1.25 起
专有 多模态
O

GPT-5.1 Codex

OpenAI

Codex GPT for repository edits, code review, and practical software agents

2025年11月13日 400K 上下文 $1.07 起
专有 多模态

Coding-optimized GPT model for repository edits, reviews, and agentic software work

2025年11月13日 400K 上下文 $1.13 起
专有 多模态

Coding-optimized GPT model for repository edits, reviews, and agentic software work

2025年11月13日 400K 上下文 $0.23 起
开源
M

Kimi K2 Thinking

Moonshot AI

Thinking Kimi model for slower research passes, planning, and hard technical questions

2025年11月6日 262K 上下文 $0.40 起
开源
M

Kimi reasoning model for long-horizon research, planning, and tool use

2025年11月6日 262K 上下文 $1.15 起
专有 多模态
A

Claude Opus 4.5

Anthropic

Flagship Claude model for deep reasoning, coding, and long-horizon agents

2025年11月1日 200K 上下文 $0.71 起

2025年10月

开源

Safety model for policy screening, moderation, and risk-aware routing workflows

2025年10月29日 131K 上下文 $0.15 起
开源 多模态

Nemotron multimodal model for visual reasoning and agentic AI workflows

2025年10月28日 128K 上下文 $0.20 起
开源
M

MiniMax-M2

MiniMax

Efficient open MiniMax model built for coding agents and tool-heavy workflows

2025年10月27日 1.0M 上下文 $0.00 起
专有 多模态
A

Claude Haiku 4.5

Anthropic

Fast Claude model for responsive assistance, classification, and lightweight agents

2025年10月15日 200K 上下文 $0.14 起
专有 多模态
A

Claude Haiku 4.5

Anthropic

Claude Haiku 4.5 is Anthropic's fastest, most cost-efficient model, matching Sonnet 4's performance on coding, computer use, and agent tasks. It offers similar performance to Sonnet 4 at one-third the cost and more than twice the speed, making it ideal for high-volume, latency-sensitive applications and multi-agent orchestration.

2025年10月15日 200K 上下文 $1.00 起
专有 多模态
O

GPT-5 Pro

OpenAI

Higher-accuracy GPT-5 tier for tough analysis, coding reviews, and planning

2025年10月6日 400K 上下文 $13.50 起

2025年9月

开源 多模态
Z

GLM-4.6

Zhipu AI

GLM-4.6 is the latest version of Z.ai's flagship model, bringing significant improvements over GLM-4.5. Key features include: 200K token context window (expanded from 128K), superior coding performance with better real-world application in Claude Code/Cline/Roo Code/Kilo Code, advanced reasoning with tool use during inference, stronger agent capabilities, and refined writing aligned with human preferences. GLM-4.6 achieves competitive performance with DeepSeek-V3.2-Exp and Claude Sonnet 4, reaching near parity with Claude Sonnet 4 (48.6% win rate) on CC-Bench real-world coding tasks.

2025年9月30日 205K 上下文 $0.00 起
专有

Speech generation model for controllable voice, narration, and audio delivery

2025年9月30日 33K 上下文 $0.50 起
专有

Speech generation model for controllable voice, narration, and audio delivery

2025年9月30日 33K 上下文 $1.00 起
专有 多模态

Balanced Claude model for coding, analysis, agent workflows, and cost control

2025年9月29日 1.0M 上下文 $3.00 起
专有 多模态
A

Claude Sonnet 4.5

Anthropic

Claude Sonnet 4.5 is the best coding model in the world. It's the strongest model for building complex agents. It’s the best model at using computers. And it shows substantial gains in reasoning and math. Highest intelligence across most tasks with exceptional agent and coding capabilities.

2025年9月29日 1.0M 上下文 $0.43 起
开源
D

DeepSeek-V3.2-Exp is an experimental iteration introducing DeepSeek Sparse Attention (DSA) to improve long-context training and inference efficiency while keeping output quality on par with V3.1. It explores fine-grained sparse attention for extended sequence processing.

2025年9月29日
专有
A

Qwen3 Max

Alibaba Cloud / Qwen Team

Flagship Qwen3 model for coding agents, complex reasoning, and tool use

2025年9月23日 262K 上下文 $0.00 起
专有 多模态
A

Qwen3-VL Plus

Alibaba Cloud / Qwen Team

Qwen vision-language model for visual reasoning, documents, and agent tasks

2025年9月23日 262K 上下文 $0.00 起
专有
O

GPT-5 Codex

OpenAI

GPT-5 Codex has been trained specifically for conducting code reviews and finding critical flaws. When reviewing, it navigates your codebase and analyzes code patterns to identify potential security vulnerabilities, performance issues, and bugs.

2025年9月15日 400K 上下文 $1.07 起
开源
A

Qwen3-Next-80B-A3B-Base

Alibaba Cloud / Qwen Team

Qwen3-Next-80B-A3B-Base is the foundation model in the Qwen3-Next series, featuring revolutionary architectural innovations for ultimate training and inference efficiency. It introduces Hybrid Attention combining Gated DeltaNet (75% layers) and Gated Attention (25% layers) for efficient ultra-long context modeling, Ultra-Sparse MoE with 512 total experts but only 10 routed + 1 shared expert activated (3.7% activation ratio), and native Multi-Token Prediction for faster inference. With 80B total parameters and only ~3B activated per inference step, it achieves performance comparable to Qwen3-32B while using less than 10% training cost and delivering 10x+ throughput for 32K+ contexts. Trained on 15T tokens with training-stability-friendly designs including Zero-Centered RMSNorm and normalized MoE router parameters. Supports 256K context length, extensible to 1M tokens with YaRN scaling.

2025年9月10日
开源
A

Qwen3-Next-80B-A3B-Instruct

Alibaba Cloud / Qwen Team

Qwen3-Next-80B-A3B-Instruct is the first in the Qwen3-Next series, featuring groundbreaking architectural innovations. It uses Hybrid Attention combining Gated DeltaNet and Gated Attention for efficient ultra-long context modeling, High-Sparsity MoE with 512 experts (10 activated + 1 shared) achieving extreme low activation ratio, and Multi-Token Prediction for improved performance and faster inference. With 80B total parameters and only 3B activated, it outperforms Qwen3-32B-Base with 10% training cost and 10x throughput for 32K+ contexts. The model performs on par with Qwen3-235B-A22B-Instruct-2507 while excelling at ultra-long-context tasks up to 256K tokens (extensible to 1M with YaRN). Architecture: 48 layers, 15T training tokens, hybrid layout of 12*(3*(Gated DeltaNet->MoE)->(Gated Attention->MoE)).

2025年9月10日 262K 上下文 $0.14 起
开源
A

Qwen3-Next-80B-A3B-Thinking

Alibaba Cloud / Qwen Team

Qwen3-Next-80B-A3B-Thinking is the thinking variant of the Qwen3-Next series, featuring the same groundbreaking architecture as the instruct model. Leveraging GSPO, it addresses stability and efficiency challenges of hybrid attention + high-sparsity MoE in RL training. It uses Hybrid Attention combining Gated DeltaNet and Gated Attention for efficient ultra-long context modeling, High-Sparsity MoE with 512 experts (10 activated + 1 shared), and Multi-Token Prediction. With 80B total parameters and only 3B activated, it demonstrates outstanding performance on complex reasoning tasks — outperforming Qwen3-30B-A3B-Thinking-2507, Qwen3-32B-Thinking, and even the proprietary Gemini-2.5-Flash-Thinking across multiple benchmarks. Architecture: 48 layers, 15T training tokens, hybrid layout of 12*(3*(Gated DeltaNet->MoE)->(Gated Attention->MoE)). Supports only thinking mode with automatic <think> tag inclusion, may generate longer thinking content.

2025年9月10日 131K 上下文 $0.14 起
专有
M

Kimi K2 0905

Moonshot AI

Kimi K2 0905 is the September update of Kimi K2 0711. It is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. It supports long-context inference up to 256k tokens, extended from the previous 128k. This update improves agentic coding with higher accuracy and better generalization across scaffolds, and enhances frontend coding with more aesthetic and functional outputs for web, 3D, and related tasks. The model is trained with a novel stack incorporating the MuonClip optimizer for stable large-scale MoE training.

2025年9月5日 262K 上下文 $0.00 起
开源
M

Kimi K2-Instruct-0905

Moonshot AI

Kimi K2-Instruct-0905 is the latest, most capable version of Kimi K2, achieving state-of-the-art performance in frontier knowledge, math, and coding among non-thinking models. This Mixture-of-Experts model features 32 billion activated parameters and 1 trillion total parameters, meticulously optimized for agentic tasks. Key features include enhanced agentic coding intelligence, extended context length to 256K tokens, and a hybrid architecture trained with MuonClip optimizer on 15.5T tokens. The model achieves 65.8% on SWE-bench Verified (single attempt), 47.3% on SWE-bench Multilingual, and excels at tool use with 70.6% on Tau2-retail. It is a reflex-grade model without long thinking, designed to act and execute complex tasks seamlessly.

2025年9月5日
开源
S

Sarvam 105B

Sarvam AI

Flagship Indian-language reasoning model for enterprise multilingual applications

2025年9月1日 131K 上下文 $0.05 起

2025年8月

专有 多模态
X

Grok 4 Fast is a high-speed variant of Grok-4, optimized for faster inference while maintaining strong reasoning capabilities. It offers improved throughput and lower latency compared to the standard Grok-4 model.

2025年8月28日
专有

Grok Code Fast 1 is a speedy and economical reasoning model that excels at agentic coding. Built from scratch with a brand-new model architecture, it features a pre-training corpus rich with programming-related content and post-training datasets that reflect real-world pull requests and coding tasks. The model has mastered the use of common tools like grep, terminal, and file editing, making it ideal for integration with IDEs. It is exceptionally versatile across the full software development stack and is particularly adept at TypeScript, Python, Java, Rust, C++, and Go.

2025年8月28日 256K 上下文 $0.18 起
开源

Translation model for multilingual conversion, localization, and cross-language workflows

2025年8月28日 8K 上下文 $2.50 起
专有 多模态
G

Nano Banana

Google

Nano Banana image model for fast generation, edits, and character-consistent assets

2025年8月26日 33K 上下文 $0.30 起
开源

Cohere reasoning model for multilingual enterprise agents, tools, and complex workflows

2025年8月21日 256K 上下文 $2.50 起
开源

NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be controlled via a system prompt. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be configured to do so, albeit with a slight decrease in accuracy for harder prompts that require reasoning. Conversely, allowing the model to generate reasoning traces first generally results in higher-quality final solutions to queries and tasks.

2025年8月18日
开源 多模态
Z

GLM-4.5V

Zhipu AI

GLM-4.5V is a multimodal (vision-language) model based on GLM-4.5-Air (106B total, 12B active) that extends hybrid reasoning to images and video. It achieves state-of-the-art results across 40+ VLM benchmarks (image reasoning, video understanding, GUI tasks, chart/document parsing, grounding) while supporting a Thinking Mode switch for deep reasoning. Released under MIT with FP8/BF16 variants and tooling in Transformers, vLLM, and SGLang.

2025年8月11日 128K 上下文 $0.29 起
专有 多模态
O

GPT-5

OpenAI

GPT-5 is our flagship model for coding, reasoning, and agentic tasks across domains. The best model for coding and agentic tasks with higher reasoning capabilities and medium speed.

2025年8月7日 400K 上下文 $1.07 起
专有 多模态

Chat-tuned GPT model for conversational assistance, writing, and tool workflows

2025年8月7日 400K 上下文 $1.13 起
专有 多模态
O

GPT-5 mini

OpenAI

A faster, more cost-efficient version of GPT-5 for well-defined tasks. Great for well-defined tasks and precise prompts with high reasoning capabilities at reduced cost.

2025年8月7日 400K 上下文 $0.04 起
专有 多模态
O

GPT-5 nano

OpenAI

GPT-5 nano is our fastest, cheapest version of GPT-5. It's great for summarization and classification tasks with average reasoning capabilities and very fast speed.

2025年8月7日 400K 上下文 $0.05 起
开源
O

GPT OSS 120B

OpenAI

GPT-OSS-120B is an open-weight, 116.8B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized to run on a single H100 GPU with native MXFP4 quantization. The model supports configurable reasoning depth, full chain-of-thought access, and native tool use, including function calling, browsing, and structured output generation. It achieves near-parity with OpenAI o4-mini on core reasoning benchmarks. Note: While referred to as '120b' for simplicity, it technically has 116.8B parameters.

2025年8月5日 131K 上下文 $0.00 起
开源
O

GPT OSS 20B

OpenAI

The gpt-oss-20b model (technically 20.9B parameters) achieves near-parity with OpenAI o4-mini on core reasoning benchmarks, while running efficiently on a single 80 GB GPU. The gpt-oss-20b model delivers similar results to OpenAI o3‑mini on common benchmarks and can run on edge devices with just 16 GB of memory, making it ideal for on-device use cases, local inference, or rapid iteration without costly infrastructure. Both models also perform strongly on tool use, few-shot function calling, CoT reasoning (as seen in results on the Tau-Bench agentic evaluation suite) and HealthBench (even outperforming proprietary models like OpenAI o1 and GPT‑4o). Note: While referred to as '20b' for simplicity, it technically has 20.9B parameters.

2025年8月5日 131K 上下文 $0.00 起
专有 多模态

Flagship Claude model for deep reasoning, coding, and long-horizon agents

2025年8月5日 200K 上下文 $15.00 起
专有 多模态
A

Claude Opus 4.1

Anthropic

Claude Opus 4.1 is a hybrid reasoning model that pushes the frontier for coding and AI agents, featuring a 200K context window. It delivers superior performance and precision for real-world coding and agentic tasks, handling complex multi-step problems with rigor and attention to detail. With extended thinking capabilities, it offers instant responses or extended step-by-step thinking visible through user-friendly summaries. It advances state-of-the-art coding performance to 74.5% on SWE-bench Verified, excels at agentic search and research, and produces human-quality content with exceptional writing abilities. It supports 32K output tokens and adapts to specific coding styles while delivering exceptional quality for extensive generation and refactoring projects.

2025年8月5日 200K 上下文 $13.50 起

2025年7月

开源 多模态
C

Cohere vision model for multilingual document analysis, OCR, and image understanding

2025年7月31日 128K 上下文 $2.50 起
专有
A

Qwen Flash

Alibaba Cloud / Qwen Team

Efficient Qwen model for fast chat, extraction, and high-volume workloads

2025年7月28日 1.0M 上下文 $0.02 起
专有
A

Qwen3 Coder Flash

Alibaba Cloud / Qwen Team

Qwen coding model for software agents, repository edits, and code reasoning

2025年7月28日 1.0M 上下文 $0.14 起
开源
Z

GLM-4.5

Zhipu AI

GLM-4.5 is an Agentic, Reasoning, and Coding (ARC) foundation model designed for intelligent agents, featuring 355 billion total parameters with 32 billion active parameters using MoE architecture. Trained on 23T tokens through multi-stage training, it is a hybrid reasoning model that provides two modes: thinking mode for complex reasoning and tool usage, and non-thinking mode for immediate responses. The model unifies agentic, reasoning, and coding capabilities with 128K context length support. It achieves exceptional performance with a score of 63.2 across 12 industry-standard benchmarks, placing 3rd among all proprietary and open-source models. Released under MIT open-source license allowing commercial use and secondary development.

2025年7月28日 131K 上下文 $0.29 起
开源
Z

GLM-4.5-Air

Zhipu AI

GLM-4.5-Air is a more compact variant of GLM-4.5 designed for efficient Agentic, Reasoning, and Coding (ARC) applications. It features 106 billion total parameters with 12 billion active parameters using MoE architecture. Like GLM-4.5, it is a hybrid reasoning model providing thinking mode for complex reasoning and tool usage, and non-thinking mode for immediate responses. Despite its compact design, GLM-4.5-Air delivers competitive performance with a score of 59.8 across 12 industry-standard benchmarks, ranking 6th overall while maintaining superior efficiency. It supports 128K context length and is released under MIT open-source license allowing commercial use.

2025年7月28日 131K 上下文 $0.00 起
专有
Z

GLM-4.5-Flash

Zhipu AI

Efficient GLM model for fast reasoning, coding, and agent workflows

2025年7月28日 131K 上下文 $0.00 起
开源
A

Qwen3-235B-A22B-Thinking-2507

Alibaba Cloud / Qwen Team

Qwen3-235B-A22B-Thinking-2507 is a state-of-the-art thinking-enabled Mixture-of-Experts (MoE) model with 235B total parameters (22B activated). It features 94 layers, 128 experts (8 activated), and supports 262K native context length. This version delivers significantly improved reasoning performance, achieving state-of-the-art results among open-source thinking models on logical reasoning, mathematics, science, coding, and academic benchmarks. Key enhancements include markedly better general capabilities (instruction following, tool usage, text generation), enhanced 256K long-context understanding, and increased thinking depth. The model supports only thinking mode with automatic <think> tag inclusion.

2025年7月25日 262K 上下文 $0.00 起
专有
A

Qwen3 Coder Plus

Alibaba Cloud / Qwen Team

Hosted Qwen coder for software agents, repo edits, and long-context code

2025年7月23日 1.0M 上下文 $0.00 起
开源
A

Qwen3-235B-A22B-Instruct-2507

Alibaba Cloud / Qwen Team

Qwen3-235B-A22B-Instruct-2507 is the updated instruct version of Qwen3-235B-A22B featuring significant improvements in general capabilities including instruction following, logical reasoning, text comprehension, mathematics, science, coding and tool usage. It provides substantial gains in long-tail knowledge coverage across multiple languages and markedly better alignment with user preferences in subjective and open-ended tasks.

2025年7月22日 262K 上下文 $0.06 起
开源
M

Devstral Small 1.1

Mistral AI

Devstral Small 1.1 (also called devstral-small-2507) is based on the Mistral-Small-3.1 foundation model and contains approximately 24 billion parameters. It supports a 128k token context window, which allows it to handle multi-file code inputs and long prompts typical in software engineering workflows. The model is fine-tuned specifically for structured outputs, including XML and function-calling formats. This makes it compatible with agent frameworks such as OpenHands and suitable for tasks like program navigation, multi-step edits, and code search. It is licensed under Apache 2.0 and available for both research and commercial use.

2025年7月11日 128K 上下文 $0.10 起
开源
M

Kimi K2 Base

Moonshot AI

Kimi K2 base model is a state-of-the-art mixture-of-experts (MoE) language model with 32 billion activated parameters and 1 trillion total parameters. Trained on 15.5 trillion tokens with the MuonClip optimizer, this is the foundation model before instruction tuning. It demonstrates strong performance on knowledge, reasoning, and coding benchmarks while being optimized for agentic capabilities.

2025年7月11日
开源
M

Kimi K2 Instruct

Moonshot AI

Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model with 32 billion activated parameters and 1 trillion total parameters. Trained with the MuonClip optimizer, it achieves exceptional performance across frontier knowledge, reasoning, and coding tasks while being meticulously optimized for agentic capabilities. The instruct variant is post-trained for drop-in, general-purpose chat and agentic experiences without long thinking.

2025年7月11日 131K 上下文 $0.55 起
专有
M

Devstral Medium

Mistral AI

Devstral Medium builds upon the strengths of Devstral Small and takes performance to the next level with a score of 61.6% on SWE-Bench Verified. Devstral Medium is available through the Mistral public API, and offers exceptional performance at a competitive price point, making it an ideal choice for businesses and developers looking for a high-quality, cost-effective model.

2025年7月10日 128K 上下文 $0.40 起
专有 多模态
X

Grok-4

xAI

Grok 4, announced by xAI in summer 2025, represents a major leap in AI capabilities, described as 'the smartest AI in the world.' Built on version 6 of xAI's foundation model, it uses 100x more training compute than Grok 2 and 10x more reinforcement learning compute than Grok 3. The model achieves PhD-level performance across all academic disciplines simultaneously, scoring perfect on standardized tests like the SAT and near-perfect on graduate exams like the GRE. Unlike Grok 3, tool usage is built into the training process rather than relying on generalization. Trained using 200,000 GPUs, Grok 4 excels at complex reasoning, mathematical problem-solving, and coding tasks, though it has acknowledged weaknesses in multimodal capabilities that are being addressed in the next version.

2025年7月9日 256K 上下文 $3.00 起
专有 多模态
X

Grok 4 Heavy is the multi-agent version of Grok 4, released alongside the standard model in summer 2025. This system spawns multiple Grok 4 agents in parallel that work independently on problems and then collaborate by comparing their solutions, similar to a study group. The agents share insights and tricks they discover, with the system intelligently combining their work rather than simply using majority voting. Grok 4 Heavy uses approximately 10x more test-time compute than regular Grok 4, enabling it to solve significantly more complex problems. On the Humanities Last Exam, it achieves over 50% accuracy on text-only problems, and it scored a perfect result on the AIME 2025 mathematics competition. The system represents a major advancement in multi-agent AI collaboration and reasoning capabilities.

2025年7月9日

2025年6月

专有 多模态
G

Gemma 3n E2B

Google

Gemma 3n is a multimodal model designed to run locally on hardware, supporting image, text, audio, and video inputs. It features a language decoder, audio encoder, and vision encoder, and is available in two sizes: E2B and E4B. The model is optimized for memory efficiency, allowing it to run on devices with limited GPU RAM. Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models. Gemma models are well-suited for a variety of content understanding tasks, including question answering, summarization, and reasoning. Their relatively small size makes it possible to deploy them in environments with limited resources such as laptops, desktops or your own cloud infrastructure, democratizing access to state of the art AI models and helping foster innovation for everyone. Gemma 3n models are designed for efficient execution on low-resource devices. They are capable of multimodal input, handling text, image, video, and audio input, and generating text outputs, with open weights for instruction-tuned variants. These models were trained with data in over 140 spoken languages.

2025年6月26日
专有 多模态

Gemma 3n is a multimodal model designed to run locally on hardware, supporting image, text, audio, and video inputs. It features a language decoder, audio encoder, and vision encoder, and is available in two sizes: E2B and E4B. The model is optimized for memory efficiency, allowing it to run on devices with limited GPU RAM. Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models. Gemma models are well-suited for a variety of content understanding tasks, including question answering, summarization, and reasoning. Their relatively small size makes it possible to deploy them in environments with limited resources such as laptops, desktops or your own cloud infrastructure, democratizing access to state of the art AI models and helping foster innovation for everyone. Gemma 3n models are designed for efficient execution on low-resource devices. They are capable of multimodal input, handling text, image, video, and audio input, and generating text outputs, with open weights for instruction-tuned variants. These models were trained with data in over 140 spoken languages.

2025年6月26日
专有 多模态
G

Gemma 3n E4B

Google

Gemma 3n is a multimodal model designed to run locally on hardware, supporting image, text, audio, and video inputs. It features a language decoder, audio encoder, and vision encoder, and is available in two sizes: E2B and E4B. The model is optimized for memory efficiency, allowing it to run on devices with limited GPU RAM. Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models. Gemma models are well-suited for a variety of content understanding tasks, including question answering, summarization, and reasoning. Their relatively small size makes it possible to deploy them in environments with limited resources such as laptops, desktops or your own cloud infrastructure, democratizing access to state of the art AI models and helping foster innovation for everyone. Gemma 3n models are designed for efficient execution on low-resource devices. They are capable of multimodal input, handling text, image, video, and audio input, and generating text outputs, with open weights for instruction-tuned variants. These models were trained with data in over 140 spoken languages.

2025年6月26日
专有 多模态

Gemma 3n is a multimodal model designed to run locally on hardware, supporting image, text, audio, and video inputs. It features a language decoder, audio encoder, and vision encoder, and is available in two sizes: E2B and E4B. The model is optimized for memory efficiency, allowing it to run on devices with limited GPU RAM. Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models. Gemma models are well-suited for a variety of content understanding tasks, including question answering, summarization, and reasoning. Their relatively small size makes it possible to deploy them in environments with limited resources such as laptops, desktops or your own cloud infrastructure, democratizing access to state of the art AI models and helping foster innovation for everyone. Gemma 3n models are designed for efficient execution on low-resource devices. They are capable of multimodal input, handling text, image, video, and audio input, and generating text outputs, with open weights for instruction-tuned variants. These models were trained with data in over 140 spoken languages.

2025年6月26日
开源 多模态
M

Mistral Small 3.2

Mistral AI

Efficient Mistral model for fast chat, extraction, and production assistants

2025年6月20日 128K 上下文 $0.10 起
开源 多模态

Mistral-Small-3.2-24B-Instruct-2506 is a minor update of Mistral-Small-3.1-24B-Instruct-2503.

2025年6月20日 131K 上下文 $0.10 起
开源 多模态

Gemini 2.5 Flash-Lite is a model developed by Google DeepMind, designed to handle various tasks including reasoning, science, mathematics, code generation, and more. It features advanced capabilities in multilingual performance and long context understanding. It is optimized for low latency use cases, supporting multimodal input with a 1 million-token context length.

2025年6月17日 1.0M 上下文 $0.09 起
开源
N

Mistral model for multilingual chat, reasoning, and tool-assisted workflows

2025年6月11日
开源 多模态
M

Magistral Medium

Mistral AI

Trained solely with reinforcement learning on top of Mistral Medium 3, Magistral Medium is a reasoning model that achieves strong performance on complex math and code tasks without relying on distillation from existing reasoning models. The training uses an RLVR framework with modifications to GRPO, enabling improved reasoning ability and multilingual consistency.

2025年6月10日 128K 上下文 $2.00 起
开源
M

Building upon Mistral Small 3.1 (2503), with added reasoning capabilities, undergoing SFT from Magistral Medium traces and RL on top, it's a small, efficient reasoning model with 24B parameters. Magistral Small can be deployed locally, fitting within a single RTX 4090 or a 32GB RAM MacBook once quantized.

2025年6月10日 33K 上下文 $0.40 起
专有 多模态
O

o3-pro

OpenAI

Version of o3 with more compute for better responses. The o3-pro model uses more compute to think harder and provide consistently better answers. Designed to tackle tough problems with advanced reasoning capabilities.

2025年6月10日 200K 上下文 $20.00 起
专有 多模态

The latest preview version of Google's most advanced reasoning Gemini model, capable of solving complex problems. Built for the agentic era with enhanced reasoning capabilities, multimodal understanding (text, image, video, audio), and a 1M token context window. Features thinking preview, code execution, grounding with Google Search, system instructions, function calling, and controlled generation. Supports up to 3,000 images per prompt, 45-60 minutes of video, and 8.4 hours of audio.

2025年6月5日 1.0M 上下文 $1.13 起

2025年5月

开源
D

DeepSeek-R1-0528

DeepSeek

DeepSeek-R1-0528 is the May 28, 2025 version of DeepSeek's reasoning model. It features advanced thinking capabilities and serves as a benchmark comparison for newer models like DeepSeek-V3.1. This model excels in complex reasoning tasks, mathematical problem-solving, and code generation through its thinking mode approach.

2025年5月28日 164K 上下文 $0.57 起
专有 多模态

Flagship Claude model for deep reasoning, coding, and long-horizon agents

2025年5月22日
专有 多模态
A

Claude Opus 4

Anthropic

Claude Opus 4 is Anthropic's most powerful model and the world's best coding model, part of the Claude 4 family. It delivers sustained performance on complex, long-running tasks and agent workflows. Opus 4 excels at coding, advanced reasoning, and can use tools (like web search) during extended thinking. It supports parallel tool execution and has improved memory capabilities.

2025年5月22日 200K 上下文 $13.50 起
专有 多模态

Balanced Claude model for coding, analysis, agent workflows, and cost control

2025年5月22日
专有 多模态
A

Claude Sonnet 4

Anthropic

Claude Sonnet 4, part of the Claude 4 family, is a significant upgrade to Claude Sonnet 3.7. It excels in coding (72.7% on SWE-bench) and reasoning, responding more precisely to instructions. Sonnet 4 offers an optimal mix of capability and practicality, with enhanced steerability, and supports extended thinking with tool use.

2025年5月22日 200K 上下文 $2.70 起
专有 多模态
G

A thinking model designed for a balance between price and performance. It builds upon Gemini 2.0 Flash with upgraded reasoning, hybrid thinking control, multimodal capabilities (text, image, video, audio input), and a 1M token input context window.

2025年5月20日 1.0M 上下文 $0.09 起
专有 多模态
G

Our most intelligent AI model, built for the agentic era. Gemini 2.5 Pro leads on common benchmarks with enhanced reasoning, multimodal capabilities (text, image, video, audio input), and a 1M token context window.

2025年5月20日 1.0M 上下文 $1.13 起
专有
G

Gemini Diffusion is a state-of-the-art, experimental text diffusion model from Google DeepMind. It explores a new kind of language model designed to provide users with greater control, creativity, and speed in text generation. Instead of predicting text token-by-token, it learns to generate outputs by refining noise step-by-step, allowing for rapid iteration and error correction during generation. Key capabilities include rapid response times (reportedly 1479 tokens/sec excluding overhead), generation of more coherent text by outputting entire blocks of tokens at once, and iterative refinement for consistent outputs. It excels at tasks like editing, including in math and code contexts.

2025年5月20日
专有

Embedding model for semantic search, retrieval, clustering, and ranking pipelines

2025年5月20日 2K 上下文 $0.15 起
开源 多模态

Gemma 3n is a generative AI model optimized for use in everyday devices, such as phones, laptops, and tablets. It features innovations like Per-Layer Embedding (PLE) parameter caching and a MatFormer model architecture for reduced compute and memory. These models handle audio, text, and visual data, though this E4B preview currently supports text and vision input. Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models, and is licensed for responsible commercial use.

2025年5月20日
开源 多模态

Gemma 3n is a generative AI model optimized for use in everyday devices, such as phones, laptops, and tablets. It features innovations like Per-Layer Embedding (PLE) parameter caching and a MatFormer model architecture for reduced compute and memory. These models handle audio, text, and visual data, though this E4B preview currently supports text and vision input. Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models, and is licensed for responsible commercial use.

2025年5月20日
开源 多模态
G

MedGemma is a collection of Gemma 3 variants that are trained for performance on medical text and image comprehension. MedGemma 4B utilizes a SigLIP image encoder that has been specifically pre-trained on a variety of de-identified medical data, including chest X-rays, dermatology images, ophthalmology images, and histopathology slides. Its LLM component is trained on a diverse set of medical data, including radiology images, histopathology patches, ophthalmology images, and dermatology images. MedGemma is a multimodal model primarily evaluated on single-image tasks. It has not been evaluated for multi-turn applications and may be more sensitive to specific prompts than its predecessor, Gemma 3. Developers should consider bias in validation data and data contamination concerns when using MedGemma.

2025年5月20日
专有 多模态
M

Mistral Medium 3

Mistral AI

Mistral model for multilingual chat, reasoning, and tool-assisted workflows

2025年5月7日 131K 上下文 $0.40 起
开源

A preliminary version of the smallest model in the upcoming Granite 4.0 family, released May 2025. It utilizes a novel hybrid Mamba-2/Transformer, fine-grained mixture of experts (MoE) architecture (7B total parameters, 1B active at inference). This preview version is partially trained (2.5T tokens) but demonstrates significant memory efficiency and performance potential, validated for at least 128K context length without positional encoding.

2025年5月2日

2025年4月

开源
M

Phi-4-mini-reasoning is designed for multi-step, logic-intensive mathematical problem-solving tasks under memory/compute constrained environments and latency bound scenarios. Some of the use cases include formal proof generation, symbolic computation, advanced word problems, and a wide range of mathematical reasoning scenarios. These models excel at maintaining context across steps, applying structured logic, and delivering accurate, reliable solutions in domains that require deep analytical thinking.

2025年4月30日 128K 上下文 $0.08 起
开源
M

Phi 4 Reasoning

Microsoft

Phi-4-reasoning is a state-of-the-art open-weight reasoning model finetuned from Phi-4 using supervised fine-tuning on a dataset of chain-of-thought traces and reinforcement learning. It focuses on math, science, and coding skills.

2025年4月30日 32K 上下文 $0.13 起
开源
M

Phi-4-reasoning-plus is a state-of-the-art open-weight reasoning model finetuned from Phi-4 using supervised fine-tuning and reinforcement learning. It focuses on math, science, and coding skills. This 'plus' version has higher accuracy due to additional RL training but may have higher latency.

2025年4月30日 32K 上下文 $0.13 起
开源
A

Qwen3 235B A22B

Alibaba Cloud / Qwen Team

Qwen3 235B A22B is a large language model developed by Alibaba, featuring a Mixture-of-Experts (MoE) architecture with 235 billion total parameters and 22 billion activated parameters. It achieves competitive results in benchmark evaluations of coding, math, general capabilities, and more, compared to other top-tier models.

2025年4月29日 131K 上下文 $0.29 起
开源
A

Qwen3 30B A3B

Alibaba Cloud / Qwen Team

Qwen3-30B-A3B is a smaller Mixture-of-Experts (MoE) model from the Qwen3 series by Alibaba, with 30.5 billion total parameters and 3.3 billion activated parameters. Features hybrid thinking/non-thinking modes, support for 119 languages, and enhanced agent capabilities. It aims to outperform previous models like QwQ-32B while using significantly fewer activated parameters.

2025年4月29日 128K 上下文 $0.08 起
开源
A

Qwen3 32B

Alibaba Cloud / Qwen Team

Qwen3-32B is a large language model from Alibaba's Qwen3 series. It features 32.8 billion parameters, a 128k token context window, support for 119 languages, and hybrid thinking modes allowing switching between deep reasoning and fast responses. It demonstrates strong performance in reasoning, instruction-following, and agent capabilities.

2025年4月29日 131K 上下文 $0.00 起
专有 多模态
O

GPT-Image-1

OpenAI

OpenAI image model for production generation, edits, and brand-safe visual workflows

2025年4月24日 $5.00 起
专有 多模态
O

o3

OpenAI

OpenAI's most powerful reasoning model. o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following. Use it to think through multi-step problems that involve analysis across text, code, and images.

2025年4月16日 200K 上下文 $2.00 起
专有 多模态
O

o4-mini

OpenAI

o4-mini is OpenAI's latest small o-series model, optimized for fast, effective reasoning with exceptionally efficient performance in coding and visual tasks. It is faster and more affordable than o3.

2025年4月16日 200K 上下文 $1.10 起
开源 多模态

Granite-3.3-8B-Base is a decoder-only language model with a 128K token context window. It improves upon Granite-3.1-8B-Base by adding support for Fill-in-the-Middle (FIM) using specialized tokens, enabling the model to generate content conditioned on both prefix and suffix. This makes it well-suited for code completion tasks

2025年4月16日
开源 多模态

Granite 3.3 models feature enhanced reasoning capabilities and support for Fill-in-the-Middle (FIM) code completion. They are built on a foundation of open-source instruction datasets with permissive licenses, alongside internally curated synthetic datasets tailored for long-context problem-solving. These models preserve the key strengths of previous Granite versions, including support for a 128K context length, strong performance in retrieval-augmented generation (RAG) and function calling, and controls for response length and originality. Granite 3.3 also delivers competitive results across general, enterprise, and safety benchmarks. Released as open source, the models are available under the Apache 2.0 license.

2025年4月16日
专有 多模态
O

GPT-4.1

OpenAI

GPT-4.1 is OpenAI's latest and most advanced flagship model, significantly improving upon GPT-4 Turbo in performance across benchmarks, speed, and cost-effectiveness.

2025年4月14日 1.0M 上下文 $2.00 起
专有 多模态
O

GPT-4.1 mini

OpenAI

GPT-4.1 mini provides a balance between intelligence, speed, and cost. It's a significant leap in small model performance, even beating GPT-4o in many benchmarks while reducing latency and cost.

2025年4月14日 1.0M 上下文 $0.40 起
专有 多模态
O

GPT-4.1 nano

OpenAI

GPT-4.1 nano is OpenAI's fastest and cheapest model available in the GPT-4.1 family. It delivers exceptional performance at a small size with its 1 million token context window. Ideal for tasks like classification or autocompletion.

2025年4月14日 1.0M 上下文 $0.08 起
开源

Flagship Nemotron model for high-throughput reasoning and complex agents

2025年4月7日 128K 上下文 $0.60 起
开源

A 253B parameter derivative of Meta Llama 3.1 405B Instruct, developed by NVIDIA using Neural Architecture Search (NAS) and vertical compression. It underwent multi-phase post-training (SFT for Math, Code, Reasoning, Chat, Tool Calling; RL with GRPO) to enhance reasoning and instruction-following. Optimized for accuracy/efficiency tradeoff on NVIDIA GPUs. Supports 128k context.

2025年4月7日
开源 多模态

Llama 4 Maverick is a natively multimodal model capable of processing both text and images. It features a 17 billion active parameter mixture-of-experts (MoE) architecture with 128 experts, supporting a wide range of multimodal tasks such as conversational interaction, image analysis, and code generation. The model includes a 1 million token context window.

2025年4月5日 1.0M 上下文 $0.12 起
开源 多模态

Open multimodal Llama for strong reasoning with efficient everyday serving

2025年4月5日 1.0M 上下文 $0.27 起
开源 多模态
M

Llama 4 Scout is a natively multimodal model capable of processing both text and images. It features a 17 billion activated parameter (109B total) mixture-of-experts (MoE) architecture with 16 experts, supporting a wide range of multimodal tasks such as conversational interaction, image analysis, and code generation. The model includes a 10 million token context window.

2025年4月5日 131K 上下文 $0.08 起
开源 多模态

Open Llama with long-context vision for efficient multimodal agents

2025年4月5日 131K 上下文 $0.18 起
开源
A

Qwen3-Coder 30B-A3B Instruct

Alibaba Cloud / Qwen Team

Smaller Qwen coder for efficient local agents and repo-level fixes

2025年4月1日 262K 上下文 $0.05 起
开源
A

Qwen3-Coder 480B-A35B Instruct

Alibaba Cloud / Qwen Team

Open Qwen coding heavyweight for repository reasoning and agentic engineering

2025年4月1日 262K 上下文 $0.30 起

2025年3月

开源 多模态
A

Qwen2.5-Omni-7B

Alibaba Cloud / Qwen Team

Qwen2.5-Omni is the flagship end-to-end multimodal model in the Qwen series. It processes diverse inputs including text, images, audio, and video, delivering real-time streaming responses through text generation and natural speech synthesis using a novel Thinker-Talker architecture.

2025年3月27日
开源
D

DeepSeek-V3 0324

DeepSeek

A powerful Mixture-of-Experts (MoE) language model with 671B total parameters (37B activated per token). Features Multi-head Latent Attention (MLA), auxiliary-loss-free load balancing, and multi-token prediction training. Pre-trained on 14.8T tokens with strong performance in reasoning, math, and code tasks.

2025年3月25日 131K 上下文 $0.25 起
开源

Llama-3.1-Nemotron-Nano-8B-v1 is a large language model (LLM) which is a derivative of Meta Llama-3.1-8B-Instruct (AKA the reference model). It is a reasoning model that is post trained for reasoning, human chat preferences, and tasks, such as RAG and tool calling.

2025年3月18日
开源

Llama-3.3-Nemotron-Super-49B-v1 is a large language model (LLM) derived from Meta Llama-3.3-70B-Instruct. It's post-trained for reasoning, chat, RAG, and tool calling, offering a balance between accuracy and efficiency (optimized for single H100). It underwent multi-phase post-training including SFT and RL (RLOO, RPO).

2025年3月18日
开源 多模态

Pretrained base model version of Mistral Small 3.1. Features improved text performance, multimodal understanding, multilingual capabilities, and an expanded 128k token context window compared to Mistral Small 3. Designed for fine-tuning.

2025年3月17日
开源 多模态

Building upon Mistral Small 3 (2501), Mistral Small 3.1 (2503) adds state-of-the-art vision understanding and enhances long context capabilities up to 128k tokens without compromising text performance. With 24 billion parameters, this model achieves top-tier capabilities in both text and vision tasks.

2025年3月17日
开源
C

Command A

Cohere

Cohere command model for multilingual enterprise agents, tools, and chat

2025年3月13日 256K 上下文 $2.50 起
开源 多模态
G

Gemma 3 12B

Google

Gemma 3 12B is a 12-billion-parameter vision-language model from Google, handling text and image input and generating text output. It features a 128K context window, multilingual support, and open weights. Suitable for question answering, summarization, reasoning, and image understanding tasks.

2025年3月12日 131K 上下文 $0.05 起
开源
G

Gemma 3 1B

Google

The Gemma 3 1B model is a lightweight, 1-billion-parameter language model by Google, optimized for efficiency on resource-limited devices. At 529MB, it processes text at 2,585 tokens/second with a context window of 128,000 tokens. It supports 35+ languages but handles text-only input, unlike larger multimodal Gemma models. This balance of speed and efficiency makes it ideal for fast text processing on mobile and low-power devices.

2025年3月12日
开源 多模态
G

Gemma 3 27B

Google

Gemma 3 27B is a 27-billion-parameter vision-language model from Google, handling text and image input and generating text output. It features a 128K context window, multilingual support, and open weights. Suitable for complex question answering, summarization, reasoning, and image understanding tasks.

2025年3月12日 40K 上下文 $0.25 起
开源 多模态
G

Gemma 3 4B

Google

Gemma 3 4B is a 4-billion-parameter vision-language model from Google, handling text and image input and generating text output. It features a 128K context window, multilingual support, and open weights. Suitable for question answering, summarization, reasoning, and image understanding tasks.

2025年3月12日
专有
A

QwQ Plus

Alibaba Cloud / Qwen Team

Qwen reasoning model for deliberate problem solving, math, and coding

2025年3月5日 131K 上下文 $0.23 起
开源
A

QwQ-32B

Alibaba Cloud / Qwen Team

A model focused on advancing AI reasoning capabilities, particularly excelling in mathematics and programming. Features deep introspection and self-questioning abilities while having some limitations in language mixing and recursive/endless reasoning patterns.

2025年3月5日 131K 上下文 $0.26 起
开源 多模态
C

Open multilingual vision model for OCR, visual reasoning, and image question answering

2025年3月4日
开源 多模态
C

Aya Vision 8B

Cohere

Compact open multilingual vision model for OCR and visual question answering

2025年3月4日

2025年2月

开源 多模态
A

Qwen2.5 VL 32B Instruct

Alibaba Cloud / Qwen Team

Qwen2.5-VL is a vision-language model from the Qwen family. Key enhancements include visual understanding (objects, text, charts, layouts), visual agent capabilities (tool use, computer/phone control), long video comprehension with event pinpointing, visual localization (bounding boxes/points), and structured output generation.

2025年2月28日
开源

Open Command R model optimized for Arabic enterprise chat, RAG, and cultural knowledge

2025年2月27日 128K 上下文 $0.04 起
专有 多模态
O

GPT-4.5

OpenAI

GPT-4.5 is OpenAI's most advanced model, offering improved reasoning, coding, and creative capabilities with faster performance and longer context handling than GPT-4. It features enhanced instruction following, reduced hallucinations, and better factual accuracy.

2025年2月27日
专有 多模态
A

Claude 3.7 Sonnet

Anthropic

The most intelligent Claude model and the first hybrid reasoning model on the market. Claude 3.7 Sonnet can produce near-instant responses or extended, step-by-step thinking that is made visible to the user. Shows particularly strong improvements in coding and front-end web development.

2025年2月24日 200K 上下文 $3.00 起
专有 多模态
X

Grok-3

xAI

Grok 3, launched by xAI on February 17, 2025, is an advanced AI model with significantly enhanced capabilities compared to Grok 2, boasting an order of magnitude increase in performance. Trained on a vast dataset that includes legal documents among others, and utilizing a massive compute infrastructure with around 200,000 GPUs in a Memphis data center, Grok 3's training used ten times more compute than its predecessor. It features specialized models like Grok 3 Reasoning and Grok 3 Mini Reasoning for complex problem-solving, and it excels in benchmarks like AIME for mathematics and GPQA for PhD-level science.

2025年2月17日 131K 上下文 $3.00 起
专有 多模态
X

Grok 3 Mini is a streamlined version of xAI's Grok 3 AI model, designed for quicker response times while maintaining utility. It's tailored for users who require speed over the comprehensive capabilities of the full Grok 3 model, making it suitable for tasks where rapid information retrieval is key. Grok 3 Mini still leverages the advanced training and data that Grok 3 was built on but offers a lighter, more efficient version for everyday use.

2025年2月17日 131K 上下文 $0.30 起
专有 多模态

A Gemini 2.0 Flash model optimized for cost efficiency and low latency

2025年2月5日 2.0M 上下文 $0.08 起
开源
M

Phi 4 Mini

Microsoft

Phi 4 Mini Instruct is a lightweight (3.8B parameters) open model built upon synthetic data and filtered web data, focusing on high-quality reasoning. It supports a 128K token context length and is enhanced for instruction adherence and safety via supervised fine-tuning and direct preference optimization.

2025年2月1日 128K 上下文 $0.08 起
开源 多模态

Phi-4-multimodal-instruct is a lightweight (5.57B parameters) open multimodal foundation model that leverages research and datasets from Phi-3.5 and 4.0. It processes text, image, and audio inputs to generate text outputs, supporting a 128K token context length. Enhanced via SFT, DPO, and RLHF for instruction following and safety.

2025年2月1日 128K 上下文 $0.07 起

2025年1月

开源 多模态

Mistral Small 3 is competitive with larger models such as Llama 3.3 70B or Qwen 32B, and is an excellent open replacement for opaque proprietary models like GPT4o-mini. Mistral Small 3 is on par with Llama 3.3 70B instruct, while being more than 3x faster on the same hardware.

2025年1月30日
开源

Mistral Small 3 is a 24B-parameter LLM licensed under Apache-2.0. It focuses on low-latency, high-efficiency instruction following, maintaining performance comparable to larger models. It provides quick, accurate responses for conversational agents, function calling, and domain-specific fine-tuning. Suitable for local inference when quantized, it rivals models 2–3× its size while using significantly fewer compute resources.

2025年1月30日
专有
O

o3-mini

OpenAI

A smaller variant of O3, expected to offer enhanced multimodal capabilities, improved reasoning, and more efficient resource utilization compared to previous models while maintaining strong performance on core tasks.

2025年1月30日 200K 上下文 $1.10 起
开源 多模态
A

Qwen2.5 VL 72B Instruct

Alibaba Cloud / Qwen Team

Qwen2.5-VL is the new flagship vision-language model of Qwen, significantly improved from Qwen2-VL. It excels at recognizing objects, analyzing text/charts/layouts in images, acting as a visual agent, understanding long videos (over 1 hour) with event pinpointing, performing visual localization (bounding boxes/points), and generating structured outputs from documents.

2025年1月26日 131K 上下文 $0.13 起
开源 多模态
A

Qwen2.5 VL 7B Instruct

Alibaba Cloud / Qwen Team

Qwen2.5-VL is a vision-language model from the Qwen family. Key enhancements include visual understanding (objects, text, charts, layouts), visual agent capabilities (tool use, computer/phone control), long video comprehension with event pinpointing, visual localization (bounding boxes/points), and structured output generation.

2025年1月26日
专有 多模态

Gemini 2.0 Flash Thinking is a enhanced reasoning model, capable of showing its thoughts to improve performance and explainability. Combining speed and performance, Gemini 2.0 Flash Thinking also excels in science and math, showing its thinking to solve complex problems.

2025年1月21日
开源
D

DeepSeek-R1

DeepSeek

DeepSeek-R1 is a reasoning-focused language model from DeepSeek that features advanced thinking capabilities. It serves as the foundation for DeepSeek's reasoning model family and pioneered their thinking mode approach for complex problem-solving tasks.

2025年1月20日 164K 上下文 $0.00 起
开源

DeepSeek-R1 is the first-generation reasoning model built atop DeepSeek-V3 (671B total parameters, 37B activated per token). It incorporates large-scale reinforcement learning (RL) to enhance its chain-of-thought and reasoning capabilities, delivering strong performance in math, code, and multi-step reasoning tasks.

2025年1月20日 131K 上下文 $0.03 起
开源

DeepSeek-R1 is the first-generation reasoning model built atop DeepSeek-V3 (671B total parameters, 37B activated per token). It incorporates large-scale reinforcement learning (RL) to enhance its chain-of-thought and reasoning capabilities, delivering strong performance in math, code, and multi-step reasoning tasks.

2025年1月20日 33K 上下文 $0.00 起
开源

DeepSeek-R1 is the first-generation reasoning model built atop DeepSeek-V3 (671B total parameters, 37B activated per token). It incorporates large-scale reinforcement learning (RL) to enhance its chain-of-thought and reasoning capabilities, delivering strong performance in math, code, and multi-step reasoning tasks.

2025年1月20日
开源

DeepSeek-R1 is the first-generation reasoning model built atop DeepSeek-V3 (671B total parameters, 37B activated per token). It incorporates large-scale reinforcement learning (RL) to enhance its chain-of-thought and reasoning capabilities, delivering strong performance in math, code, and multi-step reasoning tasks.

2025年1月20日 33K 上下文 $0.14 起
开源

DeepSeek-R1 is the first-generation reasoning model built atop DeepSeek-V3 (671B total parameters, 37B activated per token). It incorporates large-scale reinforcement learning (RL) to enhance its chain-of-thought and reasoning capabilities, delivering strong performance in math, code, and multi-step reasoning tasks.

2025年1月20日 33K 上下文 $0.29 起
开源

DeepSeek-R1 is the first-generation reasoning model built atop DeepSeek-V3 (671B total parameters, 37B activated per token). It incorporates large-scale reinforcement learning (RL) to enhance its chain-of-thought and reasoning capabilities, delivering strong performance in math, code, and multi-step reasoning tasks.

2025年1月20日 33K 上下文 $0.07 起
开源
D

DeepSeek R1 Zero

DeepSeek

DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorporates cold-start data before RL. DeepSeek-R1 achieves performance comparable to OpenAI-o1 across math, code, and reasoning tasks.

2025年1月20日
专有 多模态
M

Kimi-k1.5

Moonshot AI

Kimi 1.5 is a next-generation multimodal large language model developed by Moonshot AI. It incorporates advanced reinforcement learning (RL) and scalable multimodal reasoning, delivering state-of-the-art performance in math, code, vision, and long-context reasoning tasks.

2025年1月20日
专有 多模态
A

Qwen-Omni Turbo

Alibaba Cloud / Qwen Team

Qwen omni model for text, vision, audio, and multimodal agent tasks

2025年1月19日 33K 上下文 $0.06 起
开源
D

DeepSeek-V3.1

DeepSeek

DeepSeek-V3.1 is a hybrid model supporting both thinking and non-thinking modes through different chat templates. Built on DeepSeek-V3.1-Base with a two-phase long context extension (32K phase: 630B tokens, 128K phase: 209B tokens), it features 671B total parameters with 37B activated. Key improvements include smarter tool calling through post-training optimization, higher thinking efficiency achieving comparable quality to DeepSeek-R1-0528 while responding more quickly, and UE8M0 FP8 scale data format for model weights and activations. The model excels in both reasoning tasks (thinking mode) and practical applications (non-thinking mode), with particularly strong performance in code agent tasks, math competitions, and search-based problem solving.

2025年1月10日 131K 上下文 $0.56 起

2024年12月

开源
D

DeepSeek-V3

DeepSeek

A powerful Mixture-of-Experts (MoE) language model with 671B total parameters (37B activated per token). Features Multi-head Latent Attention (MLA), auxiliary-loss-free load balancing, and multi-token prediction training. Pre-trained on 14.8T tokens with strong performance in reasoning, math, and code tasks.

2024年12月25日 128K 上下文 $0.00 起
开源 多模态
A

QvQ-72B-Preview

Alibaba Cloud / Qwen Team

An experimental research model focusing on advanced visual reasoning and step-by-step cognitive capabilities. Achieves strong performance on multi-modal science and mathematics tasks, though exhibits some limitations such as potential language mixing and recursive reasoning loops.

2024年12月25日
专有
O

o1

OpenAI

A research preview model focused on mathematical and logical reasoning capabilities, demonstrating improved performance on tasks requiring step-by-step reasoning, mathematical problem-solving, and code generation. The model shows enhanced capabilities in formal reasoning while maintaining strong general capabilities.

2024年12月17日 200K 上下文 $15.00 起
专有 多模态
O

o1-pro

OpenAI

o1-pro is OpenAI's advanced language model optimized for complex reasoning and specialized professional tasks, offering enhanced capabilities while maintaining high efficiency.

2024年12月17日 200K 上下文 $150 起
开源 多模态
D

DeepSeek VL2

DeepSeek

An advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly improves upon its predecessor, DeepSeek-VL. DeepSeek-VL2 demonstrates superior capabilities across various tasks, including but not limited to visual question answering, optical character recognition, document/table/chart understanding, and visual grounding.

2024年12月13日
开源 多模态
D

An advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly improves upon its predecessor, DeepSeek-VL. DeepSeek-VL2 demonstrates superior capabilities across various tasks, including but not limited to visual question answering, optical character recognition, document/table/chart understanding, and visual grounding.

2024年12月13日
开源 多模态
D

An advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly improves upon its predecessor, DeepSeek-VL. DeepSeek-VL2 demonstrates superior capabilities across various tasks, including but not limited to visual question answering, optical character recognition, document/table/chart understanding, and visual grounding.

2024年12月13日
开源
M

Phi 4

Microsoft

phi-4 is a state-of-the-art open model built to excel at advanced reasoning, coding, and knowledge tasks. It leverages a blend of synthetic data, filtered web data, academic texts, and supervised fine-tuning for precision, alignment, and safety.

2024年12月12日 128K 上下文 $0.13 起
开源

Llama 3.3 is a multilingual large language model optimized for dialogue use cases across multiple languages. It is a pretrained and instruction-tuned generative model with 70 billion parameters, outperforming many open-source and closed chat models on common industry benchmarks. Llama 3.3 supports a context length of 128,000 tokens and is designed for commercial and research use in multiple languages.

2024年12月6日 131K 上下文 $0.00 起
开源
C

Command R7B

Cohere

Cohere retrieval model for long-context chat and enterprise RAG workflows

2024年12月2日 128K 上下文 $0.04 起
专有 多模态
G

Next-generation model featuring superior speed, native tool use, multimodal generation, and a 1M token context window. Supports audio, images, video, and text input with capabilities for structured outputs, function calling, code execution, search, and multimodal operations.

2024年12月1日 1.0M 上下文 $0.10 起

2024年11月

开源
A

QwQ-32B-Preview

Alibaba Cloud / Qwen Team

An experimental research model focused on advancing AI reasoning capabilities, particularly excelling in mathematics and programming. Features deep introspection and self-questioning abilities while having some limitations in language mixing and recursive reasoning patterns.

2024年11月28日
专有 多模态

GPT model for general reasoning, writing, coding, and tool-assisted tasks

2024年11月20日 128K 上下文 $2.50 起
专有 多模态
A

Nova Lite

Amazon

A low-cost multimodal model that is lightning fast for processing images, video, documents, and text.

2024年11月20日
专有
A

Nova Micro

Amazon

A text-only model that delivers lowest-latency responses at very low cost while maintaining strong performance on core language tasks. Optimized for speed and efficiency while preserving high accuracy on key benchmarks.

2024年11月20日
专有 多模态
A

Nova Pro

Amazon

Amazon Nova Pro is a highly-capable multimodal model with state-of-the-art performance across text, image, and video understanding. It excels at core capabilities like language understanding, mathematical reasoning, and multimodal tasks while offering industry-leading speed and cost efficiency.

2024年11月20日
开源
M

Mistral Large 2.1

Mistral AI

Flagship Mistral model for advanced reasoning, coding, and multilingual work

2024年11月18日 131K 上下文 $2.00 起
开源 多模态
M

Pixtral Large

Mistral AI

A 124B parameter multimodal model built on top of Mistral Large 2, featuring frontier-level image understanding capabilities. Excels at understanding documents, charts, and natural images while maintaining strong text-only performance. Features a 123B multimodal decoder and 1B parameter vision encoder with a 128K context window supporting up to 30 high-resolution images.

2024年11月18日 128K 上下文 $2.00 起
专有
A

Qwen Turbo

Alibaba Cloud / Qwen Team

Efficient Qwen model for fast chat, extraction, and high-volume workloads

2024年11月1日 1.0M 上下文 $0.04 起
开源 多模态
M

Mistral Large 3

Mistral AI

Mistral's largest general model for enterprise agents, coding, and multilingual reasoning

2024年11月1日 262K 上下文 $0.50 起
开源 多模态
M

Flagship Mistral model for advanced reasoning, coding, and multilingual work

2024年11月1日 262K 上下文 $0.50 起

2024年10月

开源
C

Open multilingual model optimized for generation across 23 languages

2024年10月24日
开源
C

Compact open multilingual model optimized for generation across 23 languages

2024年10月24日
专有
A

Claude 3.5 Haiku

Anthropic

Claude 3.5 Haiku is Anthropic's fastest model, delivering advanced coding, tool use, and reasoning capabilities at an accessible price. It excels at user-facing products, specialized sub-agent tasks, and generating personalized experiences from large data volumes. The model is particularly well-suited for code completions, interactive chatbots, data extraction, and real-time content moderation.

2024年10月22日 200K 上下文 $0.80 起
专有 多模态
A

Claude 3.5 Sonnet

Anthropic

Claude 3.5 Sonnet is a powerful AI model with industry-leading software engineering skills. It excels in coding, planning, and problem-solving, with significant improvements in agentic coding and tool use tasks. The model includes computer use capabilities in public beta, allowing it to interact with computer interfaces like a human user.

2024年10月22日
开源
M

The Ministral-8B-Instruct-2410 is an instruct fine-tuned model for local intelligence, on-device computing, and at-the-edge use cases, significantly outperforming existing models of similar size.

2024年10月16日
开源

A large language model customized by NVIDIA to improve the helpfulness of LLM generated responses. It is a fine-tuned version of Llama 3.1 70B Instruct. The model was trained using RLHF (REINFORCE) with HelpSteer2-Preference prompts.

2024年10月1日
开源 多模态
O

Open Whisper checkpoint for robust multilingual transcription and captioning

2024年10月1日 $0.00 起
开源 多模态

Speech transcription model for accurate audio-to-text and captioning workflows

2024年10月1日

2024年9月

开源 多模态

Llama 3.2 11B Vision Instruct is an instruction-tuned multimodal large language model optimized for visual recognition, image reasoning, captioning, and answering general questions about an image. It accepts text and images as input and generates text as output.

2024年9月25日 128K 上下文 $0.07 起
开源

Llama 3.2 3B Instruct is a large language model that supports a context length of 128K tokens and are state-of-the-art in their class for on-device use cases like summarization, instruction following, and rewriting tasks running locally at the edge.

2024年9月25日 33K 上下文 $0.03 起
开源 多模态

Llama 3.2 90B is a large multimodal language model optimized for visual recognition, image reasoning, and captioning tasks. It supports a context length of 128,000 tokens and is designed for deployment on edge and mobile devices, offering state-of-the-art performance in image understanding and generative tasks.

2024年9月25日
开源
A

Qwen2.5 14B Instruct

Alibaba Cloud / Qwen Team

Qwen2.5-14B-Instruct is an instruction-tuned 14.7B parameter language model, part of the Qwen2.5 series. It features significant improvements in instruction following, long text generation (8K+ tokens), structured data understanding, and JSON output generation. The model supports a 128K token context length and multilingual capabilities across 29+ languages including Chinese, English, French, Spanish, and more.

2024年9月19日
开源
A

Qwen2.5 32B Instruct

Alibaba Cloud / Qwen Team

Qwen2.5-32B-Instruct is an instruction-tuned 32 billion parameter language model, part of the Qwen2.5 series. It is designed to follow instructions, generate long texts (over 8K tokens), understand structured data (e.g., tables), and generate structured outputs, especially JSON. The model supports multilingual capabilities across over 29 languages.

2024年9月19日
开源
A

Qwen2.5 72B Instruct

Alibaba Cloud / Qwen Team

Qwen2.5-72B-Instruct is an instruction-tuned 72 billion parameter language model, part of the Qwen2.5 series. It is designed to follow instructions, generate long texts (over 8K tokens), understand structured data (e.g., tables), and generate structured outputs, especially JSON. The model supports multilingual capabilities across over 29 languages.

2024年9月19日 33K 上下文 $0.06 起
开源
A

Qwen2.5 7B Instruct

Alibaba Cloud / Qwen Team

Qwen2.5-7B-Instruct is an instruction-tuned 7B parameter language model that excels at following instructions, generating long texts (over 8K tokens), understanding structured data, and generating structured outputs like JSON. The model features enhanced capabilities in mathematics, coding, and multilingual support across 29+ languages including Chinese, English, French, Spanish, and more.

2024年9月19日
开源
A

Qwen2.5-Coder 32B Instruct

Alibaba Cloud / Qwen Team

Qwen2.5-Coder is a specialized coding model trained on 5.5 trillion tokens of code data, supporting 92 programming languages with a 128K context window. It excels in code generation, completion, repair, and multi-programming tasks while maintaining strong performance in mathematics and general capabilities.

2024年9月19日
开源
A

Qwen2.5-Coder 7B Instruct

Alibaba Cloud / Qwen Team

Qwen2.5-Coder is a specialized coding model trained on 5.5 trillion tokens of code data, supporting 92 programming languages with a 128K context window. It excels in code generation, completion, and repair while maintaining strong performance in math and general tasks. The model demonstrates exceptional capabilities in multi-programming language tasks and code reasoning.

2024年9月19日
开源
M

Mistral Small

Mistral AI

An enterprise-grade 22B parameter model optimized for tasks like translation, summarization, and sentiment analysis. Offers significant improvements in human alignment, reasoning capabilities, and code generation compared to previous versions.

2024年9月17日 256K 上下文 $0.15 起
开源 多模态
M

Pixtral-12B

Mistral AI

A 12B parameter multimodal model with a 400M parameter vision encoder, capable of understanding both natural images and documents. Excels at multimodal tasks while maintaining strong text-only performance. Supports variable image sizes and multiple images in context.

2024年9月17日 128K 上下文 $0.15 起
专有
O

o1-mini

OpenAI

o1-mini is a cost-efficient language model developed by OpenAI, designed for advanced reasoning tasks while minimizing computational resources.

2024年9月12日 128K 上下文 $1.10 起
专有
O

o1-preview

OpenAI

A research preview model focused on mathematical and logical reasoning capabilities, demonstrating improved performance on tasks requiring step-by-step reasoning, mathematical problem-solving, and code generation. The model shows enhanced capabilities in formal reasoning while maintaining strong general capabilities.

2024年9月12日

2024年8月

开源
C

Command R+

Cohere

C4AI Command R+ is a 104 billion parameter model with advanced capabilities, including Retrieval Augmented Generation (RAG) and multi-step tool use, optimized for multilingual tasks.

2024年8月30日 128K 上下文 $0.15 起
开源
C

Command R+

Cohere

Cohere's RAG workhorse for long-context enterprise search and tool use

2024年8月30日 128K 上下文 $2.50 起
开源 多模态
A

Qwen2-VL-72B-Instruct

Alibaba Cloud / Qwen Team

An instruction-tuned, large multimodal model that excels at visual understanding and step-by-step reasoning. It supports image and video input, with dynamic resolution handling and improved positional embeddings (M-ROPE), enabling advanced capabilities such as complex problem solving, multilingual text recognition in images, and agent-like interactions in video contexts.

2024年8月29日
开源
M

Phi-3.5-mini-instruct is a 3.8B-parameter model that supports up to 128K context tokens, with improved multilingual capabilities across over 20 languages. It underwent additional training and safety post-training to enhance instruction-following, reasoning, math, and code generation. Ideal for environments with memory or latency constraints, it uses an MIT license.

2024年8月23日 128K 上下文 $0.13 起
开源
M

Phi-3.5-MoE-instruct is a mixture-of-experts model with ~42B total parameters (6.6B active) and a 128K context window. It excels at reasoning, math, coding, and multilingual tasks, outperforming larger dense models in many benchmarks. It underwent a thorough safety post-training process (SFT + DPO) and is licensed under MIT. This model is ideal for scenarios where efficiency and high performance are both required, particularly in multi-lingual or reasoning-intensive tasks.

2024年8月23日 128K 上下文 $0.16 起
开源 多模态

Phi-3.5-vision-instruct is a 4.2B-parameter open multimodal model with up to 128K context tokens. It emphasizes multi-frame image understanding and reasoning, boosting performance on single-image benchmarks while enabling multi-image comparison, summarization, and even video analysis. The model underwent safety post-training for improved instruction-following, alignment, and robust handling of visual and text inputs, and is released under the MIT license.

2024年8月23日
开源
A

Jamba 1.5 Large

AI21 Labs

State-of-the-art hybrid SSM-Transformer instruction following foundation model, offering superior long context handling, speed, and quality.

2024年8月22日
开源
A

Jamba 1.5 Mini

AI21 Labs

Part of the Jamba 1.5 family, a state-of-the-art hybrid SSM-Transformer instruction following foundation model offering superior long context handling, speed, and quality.

2024年8月22日
开源

Compact Nemotron model for efficient reasoning and deployable AI agents

2024年8月21日
专有 多模态
X

Grok-2

xAI

Grok-2 is a frontier language model with state-of-the-art reasoning capabilities, featuring advanced abilities in chat, coding, and reasoning. It demonstrates superior performance in visual math reasoning, document-based question answering, and excels across various academic benchmarks including reasoning, reading comprehension, math, and science.

2024年8月13日
专有 多模态
X

Grok-2 mini is a smaller, faster variant of Grok-2 that offers a balance between speed and answer quality. While more compact than its larger sibling, it maintains strong capabilities across various tasks including reasoning, coding, and chat interactions.

2024年8月13日
专有 多模态
O

GPT-4o

OpenAI

GPT-4o ('o' for 'omni') is a multimodal AI model that accepts text, audio, image, and video inputs, and generates text, audio, and image outputs. It matches GPT-4 Turbo performance on text and code, with improvements in non-English languages, vision, and audio understanding.

2024年8月6日 128K 上下文 $2.50 起

2024年7月

开源
M

Mistral Large 2

Mistral AI

A 123B parameter model with strong capabilities in code generation, mathematics, and reasoning. Features enhanced multilingual support across dozens of languages, 128k context window, and advanced function calling capabilities. Excels in instruction-following and maintains concise outputs.

2024年7月24日
开源

Llama 3.1 405B Instruct is a large language model optimized for multilingual dialogue use cases. It outperforms many available open source and closed chat models on common industry benchmarks. The model supports 8 languages and has a 128K token context length.

2024年7月23日 128K 上下文 $0.00 起
开源

Llama 3.1 70B Instruct is a large language model optimized for multilingual dialogue use cases. It outperforms many available open source and closed chat models on common industry benchmarks.

2024年7月23日 128K 上下文 $0.72 起
开源

Llama 3.1 8B Instruct is a multilingual large language model optimized for dialogue use cases. It features a 128K context length, state-of-the-art tool use, and strong reasoning capabilities.

2024年7月23日 120K 上下文 $0.02 起
开源
A

Qwen2 72B Instruct

Alibaba Cloud / Qwen Team

Qwen2-72B-Instruct is an instruction-tuned language model with 72 billion parameters, supporting a context length of up to 131,072 tokens. It's part of the new Qwen2 series, which has surpassed most open-source models and demonstrates competitiveness against proprietary models across various benchmarks.

2024年7月23日
开源
A

Qwen2 7B Instruct

Alibaba Cloud / Qwen Team

Qwen2-7B-Instruct is an instruction-tuned language model with 7 billion parameters, supporting a context length of up to 131,072 tokens.

2024年7月23日
专有 多模态
O

GPT-4o mini

OpenAI

GPT-4o mini is OpenAI's latest cost-efficient small model, designed to make AI intelligence more accessible and affordable. It excels in textual intelligence and multimodal reasoning, outperforming previous models like GPT-3.5 Turbo. With a context window of 128K tokens and support for text and vision, it offers low-cost, real-time applications such as customer support chatbots. Priced at 15 cents per million input tokens and 60 cents per million output tokens, it is significantly cheaper than its predecessors. Safety is prioritized with built-in measures and improved resistance to security threats.

2024年7月18日 128K 上下文 $0.15 起
开源
M

A state-of-the-art 12B multilingual model with a 128k context window, designed for global applications and strong in multiple languages.

2024年7月18日 128K 上下文 $0.14 起
开源
M

Mistral Nemo

Mistral AI

Efficient Mistral-NVIDIA open model for multilingual chat and local deployment

2024年7月1日 128K 上下文 $0.15 起

2024年6月

开源
G

Gemma 2 27B

Google

Gemma 2 27B IT is an instruction-tuned version of Google's state-of-the-art open language model. Built from the same research and technology as Gemini, it's optimized for dialogue applications through supervised fine-tuning, distillation from larger models, and RLHF. The model excels at text generation tasks including question answering, summarization, and reasoning.

2024年6月27日
开源
G

Gemma 2 9B

Google

Gemma 2 9B IT is an instruction-tuned version of Google's Gemma 2 9B base model. It was trained on 8 trillion tokens of web data, code, and math content. The model features sliding window attention, logit soft-capping, and knowledge distillation techniques. It's optimized for dialogue applications through supervised fine-tuning, distillation, RLHF, and model merging using WARP.

2024年6月27日
专有 多模态
O

Research model for long-horizon investigation, synthesis, and analytical reports

2024年6月26日 200K 上下文 $10.00 起
专有 多模态

Research model for long-horizon investigation, synthesis, and analytical reports

2024年6月26日 200K 上下文 $2.00 起
专有 多模态
A

Claude 3.5 Sonnet

Anthropic

Claude 3.5 Sonnet is a powerful AI model. It excels in graduate-level reasoning, undergraduate-level knowledge, and coding proficiency, with improved understanding of nuance, humor, and complex instructions.

2024年6月21日

2024年5月

开源
M

Codestral-22B

Mistral AI

A 22B parameter code generation model trained on 80+ programming languages including Python, Java, C, C++, JavaScript, and Bash. Supports both instruction-following and fill-in-the-middle (FIM) capabilities for code completion and generation tasks.

2024年5月29日
开源
M

Codestral (latest)

Mistral AI

Mistral code model for completions, refactors, and developer IDE workflows

2024年5月29日 256K 上下文 $0.30 起
专有 多模态
O

GPT-4o

OpenAI

Omni-era GPT for multimodal chat, practical coding, and general assistants

2024年5月13日 128K 上下文 $2.50 起
专有 多模态
O

GPT-4o

OpenAI

GPT-4o ('o' for 'omni') is a multimodal AI model that accepts text, audio, image, and video inputs, and generates text, audio, and image outputs. It matches GPT-4 Turbo performance on text and code, with improvements in non-English languages, vision, and audio understanding.

2024年5月13日 128K 上下文 $5.00 起
开源
D

DeepSeek-V2.5

DeepSeek

DeepSeek-V2.5 is an upgraded version that combines DeepSeek-V2-Chat and DeepSeek-Coder-V2-Instruct, integrating general and coding abilities. It better aligns with human preferences and has been optimized in various aspects, including writing and instruction following.

2024年5月8日
专有 多模态
G

Gemini 1.5 Flash is a fast and versatile multimodal model for scaling across diverse tasks. It supports audio, images, video, and text input, and produces text output. The model is optimized for generating code, extracting data, editing text, and more, making it ideal for narrow, high-frequency tasks.

2024年5月1日
专有 多模态
G

Gemini 1.5 Pro is a mid-size multimodal model optimized for a wide range of reasoning tasks. It can process large amounts of data at once, including 2 hours of video, 19 hours of audio, codebases with 60,000 lines of code, or 2,000 pages of text.

2024年5月1日

2024年4月

专有 多模态
X

Grok-1.5V

xAI

A multimodal model capable of processing text and visual information, including documents, diagrams, charts, screenshots, and photographs. Notable for strong real-world spatial understanding capabilities.

2024年4月12日
专有
O

GPT-4 Turbo

OpenAI

The latest GPT-4 model with improved performance, updated knowledge, and enhanced capabilities. It offers faster response times and more affordable pricing compared to previous versions.

2024年4月9日 128K 上下文 $10.00 起
专有 多模态
A

Qwen-VL Max

Alibaba Cloud / Qwen Team

Qwen vision-language model for visual reasoning, documents, and agent tasks

2024年4月8日 131K 上下文 $0.23 起
专有
A

Qwen Max

Alibaba Cloud / Qwen Team

Flagship Qwen model for complex reasoning, coding, and agentic workflows

2024年4月3日 131K 上下文 $0.35 起

2024年3月

专有
X

Grok-1.5

xAI

An advanced language model with improved reasoning capabilities, particularly excelling in coding and mathematical tasks. Features a 128K token context window and enhanced problem-solving abilities compared to its predecessor.

2024年3月28日
专有 多模态

A multimodal model capable of processing audio, images, video, and text with high efficiency. Features JSON mode, function calling, code execution, and system instructions support. Optimized for fast inference with 8B parameters.

2024年3月15日
专有 多模态
A

Claude 3 Haiku

Anthropic

Claude 3 Haiku is the fastest and most compact model in the Claude 3 family, designed for near-instant responsiveness. It excels at answering simple queries and requests with unmatched speed, making it ideal for seamless AI experiences that mimic human interactions.

2024年3月13日 200K 上下文 $0.25 起

2024年2月

专有 多模态
A

Claude 3 Opus

Anthropic

Claude 3 Opus is Anthropic's most intelligent model, with best-in-market performance on highly complex tasks. It can navigate open-ended prompts and sight-unseen scenarios with remarkable fluency and human-like understanding, showing the outer limits of what's possible with generative AI.

2024年2月29日
专有 多模态
A

Claude 3 Sonnet

Anthropic

Claude 3 Sonnet strikes the ideal balance between intelligence and speed—particularly for enterprise workloads. It delivers strong performance at a lower cost compared to its peers, and is engineered for high endurance in large-scale AI deployments.

2024年2月29日
专有
G

Gemini 1.0 Pro is a Natural Language Processing (NLP) model designed for tasks such as multi-turn text and code chat, and code generation. It supports text input and output, making it ideal for natural language tasks. The model is optimized for handling complex conversations and generating code snippets. It offers adjustable safety settings and supports function calling, but does not support JSON mode, JSON schema, or system instructions. The latest stable version is gemini-1.0-pro-001, and it was last updated in February 2024.

2024年2月15日

2024年1月

专有
A

Qwen Plus

Alibaba Cloud / Qwen Team

Qwen instruction model for multilingual chat, reasoning, and tool use

2024年1月25日 1.0M 上下文 $0.12 起
专有 多模态
A

Qwen-VL Plus

Alibaba Cloud / Qwen Team

Qwen vision-language model for visual reasoning, documents, and agent tasks

2024年1月25日 131K 上下文 $0.12 起
专有
P

Sonar

Perplexity

Fast web-grounded Sonar for current answers, citations, and lightweight retrieval

2024年1月1日 130K 上下文 $1.00 起
专有 多模态
P

Sonar Pro

Perplexity

Deeper Sonar search model with broader retrieval and stronger synthesis

2024年1月1日 200K 上下文 $2.99 起
专有 多模态
P

Sonar Reasoning Pro

Perplexity

Web-grounded Sonar for multi-step research questions that need cited reasoning

2024年1月1日 128K 上下文 $2.00 起

2023年6月

专有 多模态
O

GPT-4

OpenAI

GPT-4 is a large multimodal model capable of processing both image and text inputs and generating human-like text outputs. It demonstrates human-level performance on various professional and academic benchmarks.

2023年6月13日 8K 上下文 $30.00 起

2023年3月

专有
O

GPT-3.5 Turbo

OpenAI

The latest GPT-3.5 Turbo model with higher accuracy at responding in requested formats and a fix for a bug which caused a text encoding issue for non-English language function calls.

2023年3月21日 16K 上下文 $0.50 起