We are thrilled to unveil the Gemma 4 12B, our newest offering designed to integrate agentic multimodal intelligence seamlessly into laptops. This innovative model serves as a bridge between our edge-friendly E4B and the more sophisticated 26B Mixture of Experts (MoE), providing an impressive array of features within a compact memory footprint. Moreover, it marks our first mid-sized model equipped with native audio inputs.
Thanks to the contributions of the developer community, Gemma 4 models have achieved over 150 million downloads. From creating wearable robotic arms for personal assistance to high-level AI security solutions for businesses, the innovations you’ve introduced using our technology have been remarkable. We’re eager to see what new projects you will undertake with this addition to the lineup.
Here’s a glimpse of what differentiates Gemma 4 12B:
- **Innovative Unified Architecture**: This model eliminates the need for multimodal encoders, allowing vision and audio inputs to connect directly to the LLM backbone. - **Enhanced Reasoning Capabilities**: It offers benchmark performance that approaches our 26B model, allowing for sophisticated multi-step reasoning and agent-based workflows.
- **Optimized for Laptops**: The Gemma 4 12B is designed to function efficiently on local devices with just 16GB of VRAM or unified memory.
- **Open Source Support**: Released under the Apache 2.0 license, it is fully supported across the developer community, promoting accessibility and collaboration.
- **Latency Reduction Tools**: This model includes Multi-Token Prediction (MTP) drafters, designed to enhance performance by decreasing latency.
These features collectively empower users to harness advanced multimodal capabilities on everyday hardware, all without compromising on speed or reasoning depth. Let’s delve deeper into how Gemma 4 12B makes this possible.
Experience cutting-edge agents directly on your device
Gemma 4 12B offers performance metrics that closely rival those of our larger 26B MoE model while keeping the total memory consumption below half that of its predecessor. Small enough to operate on standard consumer laptops with 16GB of RAM, it enables users to enjoy powerful multimodal and agentic experiences directly on their devices.

