Organization

Google

US 49 个模型

Technology giant with AI research

旗下模型

专有
G

Gemini 1.0 Pro is a Natural Language Processing (NLP) model designed for tasks such as multi-turn text and code chat, and code generation. It supports text input and output, making it ideal for natural language tasks. The model is optimized for handling complex conversations and generating code snippets. It offers adjustable safety settings and supports function calling, but does not support JSON mode, JSON schema, or system instructions. The latest stable version is gemini-1.0-pro-001, and it was last updated in February 2024.

2024年2月15日
专有 多模态
G

Gemini 1.5 Flash is a fast and versatile multimodal model for scaling across diverse tasks. It supports audio, images, video, and text input, and produces text output. The model is optimized for generating code, extracting data, editing text, and more, making it ideal for narrow, high-frequency tasks.

2024年5月1日
专有 多模态

A multimodal model capable of processing audio, images, video, and text with high efficiency. Features JSON mode, function calling, code execution, and system instructions support. Optimized for fast inference with 8B parameters.

2024年3月15日
专有 多模态
G

Gemini 1.5 Pro is a mid-size multimodal model optimized for a wide range of reasoning tasks. It can process large amounts of data at once, including 2 hours of video, 19 hours of audio, codebases with 60,000 lines of code, or 2,000 pages of text.

2024年5月1日
专有 多模态
G

Next-generation model featuring superior speed, native tool use, multimodal generation, and a 1M token context window. Supports audio, images, video, and text input with capabilities for structured outputs, function calling, code execution, search, and multimodal operations.

2024年12月1日 1.0M 上下文 $0.10 起
专有 多模态

Gemini 2.0 Flash Thinking is a enhanced reasoning model, capable of showing its thoughts to improve performance and explainability. Combining speed and performance, Gemini 2.0 Flash Thinking also excels in science and math, showing its thinking to solve complex problems.

2025年1月21日
专有 多模态

A Gemini 2.0 Flash model optimized for cost efficiency and low latency

2025年2月5日 2.0M 上下文 $0.08 起
专有 多模态
G

A thinking model designed for a balance between price and performance. It builds upon Gemini 2.0 Flash with upgraded reasoning, hybrid thinking control, multimodal capabilities (text, image, video, audio input), and a 1M token input context window.

2025年5月20日 1.0M 上下文 $0.09 起
专有

Speech generation model for controllable voice, narration, and audio delivery

2025年9月30日 33K 上下文 $0.50 起
开源 多模态

Gemini 2.5 Flash-Lite is a model developed by Google DeepMind, designed to handle various tasks including reasoning, science, mathematics, code generation, and more. It features advanced capabilities in multilingual performance and long context understanding. It is optimized for low latency use cases, supporting multimodal input with a 1 million-token context length.

2025年6月17日 1.0M 上下文 $0.09 起
专有 多模态
G

Our most intelligent AI model, built for the agentic era. Gemini 2.5 Pro leads on common benchmarks with enhanced reasoning, multimodal capabilities (text, image, video, audio input), and a 1M token context window.

2025年5月20日 1.0M 上下文 $1.13 起
专有 多模态

The latest preview version of Google's most advanced reasoning Gemini model, capable of solving complex problems. Built for the agentic era with enhanced reasoning capabilities, multimodal understanding (text, image, video, audio), and a 1M token context window. Features thinking preview, code execution, grounding with Google Search, system instructions, function calling, and controlled generation. Supports up to 3,000 images per prompt, 45-60 minutes of video, and 8.4 hours of audio.

2025年6月5日 1.0M 上下文 $1.13 起
专有

Speech generation model for controllable voice, narration, and audio delivery

2025年9月30日 33K 上下文 $1.00 起
专有 多模态

New Gemini flash lane bringing frontier-style multimodal reasoning to cheaper runs

2025年12月17日 1.0M 上下文 $0.07 起
专有 多模态

Preview Gemini flagship for complex reasoning, coding, and rich multimodal prompts

2025年11月18日 1.0M 上下文 $0.57 起
专有 多模态

Low-latency Gemini model for high-volume multimodal and agent workloads

2026年5月7日 1.0M 上下文 $0.25 起
专有 多模态

Low-latency Gemini model for high-volume multimodal and agent workloads

2026年3月3日 1.0M 上下文 $0.25 起
专有 多模态

Reasoning-first Gemini preview for agentic coding and complex problem solving

2026年2月19日 1.0M 上下文 $2.00 起
专有 多模态

Advanced Gemini model for complex reasoning, coding, and multimodal analysis

2026年2月19日 1.0M 上下文 $2.00 起
专有 多模态
G

Fast Gemini model balancing multimodal reasoning, tool use, and cost

2026年5月19日 1.0M 上下文 $0.19 起
新发布 专有 多模态

Fast Gemini model balancing multimodal reasoning, tool use, and cost

2026年7月21日 1.0M 上下文 $0.30 起
新发布 专有 多模态
G

Fast Gemini model balancing multimodal reasoning, tool use, and cost

2026年7月21日 1.0M 上下文 $1.50 起
专有
G

Gemini Diffusion is a state-of-the-art, experimental text diffusion model from Google DeepMind. It explores a new kind of language model designed to provide users with greater control, creativity, and speed in text generation. Instead of predicting text token-by-token, it learns to generate outputs by refining noise step-by-step, allowing for rapid iteration and error correction during generation. Key capabilities include rapid response times (reportedly 1479 tokens/sec excluding overhead), generation of more coherent text by outputting entire blocks of tokens at once, and iterative refinement for consistent outputs. It excels at tasks like editing, including in math and code contexts.

2025年5月20日
专有

Embedding model for semantic search, retrieval, clustering, and ranking pipelines

2025年5月20日 2K 上下文 $0.15 起
专有 多模态

Fast Gemini model balancing multimodal reasoning, tool use, and cost

2026年5月19日 1.0M 上下文 $1.50 起
专有 多模态

Low-latency Gemini model for high-volume multimodal and agent workloads

2026年5月7日 1.0M 上下文 $0.25 起
新发布 专有 多模态

Video generation and editing model for fast, conversational text- and image-to-video workflows

2026年6月30日 131K 上下文 $1.50 起
开源
G

Gemma 2 27B

Google

Gemma 2 27B IT is an instruction-tuned version of Google's state-of-the-art open language model. Built from the same research and technology as Gemini, it's optimized for dialogue applications through supervised fine-tuning, distillation from larger models, and RLHF. The model excels at text generation tasks including question answering, summarization, and reasoning.

2024年6月27日
开源
G

Gemma 2 9B

Google

Gemma 2 9B IT is an instruction-tuned version of Google's Gemma 2 9B base model. It was trained on 8 trillion tokens of web data, code, and math content. The model features sliding window attention, logit soft-capping, and knowledge distillation techniques. It's optimized for dialogue applications through supervised fine-tuning, distillation, RLHF, and model merging using WARP.

2024年6月27日
开源 多模态
G

Gemma 3 12B

Google

Gemma 3 12B is a 12-billion-parameter vision-language model from Google, handling text and image input and generating text output. It features a 128K context window, multilingual support, and open weights. Suitable for question answering, summarization, reasoning, and image understanding tasks.

2025年3月12日 131K 上下文 $0.05 起
开源
G

Gemma 3 1B

Google

The Gemma 3 1B model is a lightweight, 1-billion-parameter language model by Google, optimized for efficiency on resource-limited devices. At 529MB, it processes text at 2,585 tokens/second with a context window of 128,000 tokens. It supports 35+ languages but handles text-only input, unlike larger multimodal Gemma models. This balance of speed and efficiency makes it ideal for fast text processing on mobile and low-power devices.

2025年3月12日
开源 多模态
G

Gemma 3 27B

Google

Gemma 3 27B is a 27-billion-parameter vision-language model from Google, handling text and image input and generating text output. It features a 128K context window, multilingual support, and open weights. Suitable for complex question answering, summarization, reasoning, and image understanding tasks.

2025年3月12日 40K 上下文 $0.25 起
开源 多模态
G

Gemma 3 4B

Google

Gemma 3 4B is a 4-billion-parameter vision-language model from Google, handling text and image input and generating text output. It features a 128K context window, multilingual support, and open weights. Suitable for question answering, summarization, reasoning, and image understanding tasks.

2025年3月12日
专有 多模态
G

Gemma 3n E2B

Google

Gemma 3n is a multimodal model designed to run locally on hardware, supporting image, text, audio, and video inputs. It features a language decoder, audio encoder, and vision encoder, and is available in two sizes: E2B and E4B. The model is optimized for memory efficiency, allowing it to run on devices with limited GPU RAM. Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models. Gemma models are well-suited for a variety of content understanding tasks, including question answering, summarization, and reasoning. Their relatively small size makes it possible to deploy them in environments with limited resources such as laptops, desktops or your own cloud infrastructure, democratizing access to state of the art AI models and helping foster innovation for everyone. Gemma 3n models are designed for efficient execution on low-resource devices. They are capable of multimodal input, handling text, image, video, and audio input, and generating text outputs, with open weights for instruction-tuned variants. These models were trained with data in over 140 spoken languages.

2025年6月26日
专有 多模态

Gemma 3n is a multimodal model designed to run locally on hardware, supporting image, text, audio, and video inputs. It features a language decoder, audio encoder, and vision encoder, and is available in two sizes: E2B and E4B. The model is optimized for memory efficiency, allowing it to run on devices with limited GPU RAM. Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models. Gemma models are well-suited for a variety of content understanding tasks, including question answering, summarization, and reasoning. Their relatively small size makes it possible to deploy them in environments with limited resources such as laptops, desktops or your own cloud infrastructure, democratizing access to state of the art AI models and helping foster innovation for everyone. Gemma 3n models are designed for efficient execution on low-resource devices. They are capable of multimodal input, handling text, image, video, and audio input, and generating text outputs, with open weights for instruction-tuned variants. These models were trained with data in over 140 spoken languages.

2025年6月26日
开源 多模态

Gemma 3n is a generative AI model optimized for use in everyday devices, such as phones, laptops, and tablets. It features innovations like Per-Layer Embedding (PLE) parameter caching and a MatFormer model architecture for reduced compute and memory. These models handle audio, text, and visual data, though this E4B preview currently supports text and vision input. Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models, and is licensed for responsible commercial use.

2025年5月20日
专有 多模态
G

Gemma 3n E4B

Google

Gemma 3n is a multimodal model designed to run locally on hardware, supporting image, text, audio, and video inputs. It features a language decoder, audio encoder, and vision encoder, and is available in two sizes: E2B and E4B. The model is optimized for memory efficiency, allowing it to run on devices with limited GPU RAM. Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models. Gemma models are well-suited for a variety of content understanding tasks, including question answering, summarization, and reasoning. Their relatively small size makes it possible to deploy them in environments with limited resources such as laptops, desktops or your own cloud infrastructure, democratizing access to state of the art AI models and helping foster innovation for everyone. Gemma 3n models are designed for efficient execution on low-resource devices. They are capable of multimodal input, handling text, image, video, and audio input, and generating text outputs, with open weights for instruction-tuned variants. These models were trained with data in over 140 spoken languages.

2025年6月26日
专有 多模态

Gemma 3n is a multimodal model designed to run locally on hardware, supporting image, text, audio, and video inputs. It features a language decoder, audio encoder, and vision encoder, and is available in two sizes: E2B and E4B. The model is optimized for memory efficiency, allowing it to run on devices with limited GPU RAM. Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models. Gemma models are well-suited for a variety of content understanding tasks, including question answering, summarization, and reasoning. Their relatively small size makes it possible to deploy them in environments with limited resources such as laptops, desktops or your own cloud infrastructure, democratizing access to state of the art AI models and helping foster innovation for everyone. Gemma 3n models are designed for efficient execution on low-resource devices. They are capable of multimodal input, handling text, image, video, and audio input, and generating text outputs, with open weights for instruction-tuned variants. These models were trained with data in over 140 spoken languages.

2025年6月26日
开源 多模态

Gemma 3n is a generative AI model optimized for use in everyday devices, such as phones, laptops, and tablets. It features innovations like Per-Layer Embedding (PLE) parameter caching and a MatFormer model architecture for reduced compute and memory. These models handle audio, text, and visual data, though this E4B preview currently supports text and vision input. Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models, and is licensed for responsible commercial use.

2025年5月20日
开源 多模态

Open Gemma instruction model for efficient chat and self-hosted deployments

2026年4月2日 262K 上下文 $0.07 起
开源 多模态
G

Largest Gemma 4 instruction model for open, self-hosted chat and reasoning

2026年4月2日 262K 上下文 $0.00 起
开源 多模态
G

Open Gemma instruction model for efficient chat and self-hosted deployments

2026年4月2日
开源 多模态
G

Open Gemma instruction model for efficient chat and self-hosted deployments

2026年4月2日
开源 多模态
G

MedGemma is a collection of Gemma 3 variants that are trained for performance on medical text and image comprehension. MedGemma 4B utilizes a SigLIP image encoder that has been specifically pre-trained on a variety of de-identified medical data, including chest X-rays, dermatology images, ophthalmology images, and histopathology slides. Its LLM component is trained on a diverse set of medical data, including radiology images, histopathology patches, ophthalmology images, and dermatology images. MedGemma is a multimodal model primarily evaluated on single-image tasks. It has not been evaluated for multi-turn applications and may be more sensitive to specific prompts than its predecessor, Gemma 3. Developers should consider bias in validation data and data contamination concerns when using MedGemma.

2025年5月20日
专有 多模态
G

Nano Banana

Google

Nano Banana image model for fast generation, edits, and character-consistent assets

2025年8月26日 33K 上下文 $0.30 起
专有 多模态
G

Nano Banana 2

Google

Image model for prompt-driven generation, editing, and visual design workflows

2026年5月28日 131K 上下文 $0.50 起
专有 多模态
G

Nano Banana 2

Google

Image model for prompt-driven generation, editing, and visual design workflows

2026年2月26日 1.0M 上下文 $0.50 起
专有 多模态
G

Nano Banana Pro for higher-fidelity image generation and design-heavy edits

2026年5月28日 66K 上下文 $2.00 起
专有 多模态
G

Nano Banana Pro for higher-fidelity image generation and design-heavy edits

2025年11月20日 1.0M 上下文 $2.00 起