Model

Gemma 3n E2B

Google
专有 Proprietary 多模态

Gemma 3n is a multimodal model designed to run locally on hardware, supporting image, text, audio, and video inputs. It features a language decoder, audio encoder, and vision encoder, and is available in two sizes: E2B and E4B. The model is optimized for memory efficiency, allowing it to run on devices with limited GPU RAM. Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models. Gemma models are well-suited for a variety of content understanding tasks, including question answering, summarization, and reasoning. Their relatively small size makes it possible to deploy them in environments with limited resources such as laptops, desktops or your own cloud infrastructure, democratizing access to state of the art AI models and helping foster innovation for everyone. Gemma 3n models are designed for efficient execution on low-resource devices. They are capable of multimodal input, handling text, image, video, and audio input, and generating text outputs, with open weights for instruction-tuned variants. These models were trained with data in over 140 spoken languages.

发布日期2025年6月26日
参数规模8B
上下文长度
许可证Proprietary
知识截止2024年6月1日

Benchmarks

评测成绩

评测基准 类别 分数 来源
PIQA reasoning physics general 78.9 来源
BoolQ language reasoning 76.4 来源
ARC-E reasoning general 75.8 来源
HellaSwag reasoning 72.2 来源
Winogrande reasoning language 66.8 来源
TriviaQA general reasoning 60.8 来源
DROP reasoning math 53.9 来源
ARC-C reasoning general 51.7 来源
Social IQa reasoning psychology 48.8 来源
BIG-Bench Hard reasoning math language 44.3 来源
Natural Questions reasoning general search 15.5 来源

Pricing

API 价格对比

暂无 API 价格。