Developers and users involved in creating production AI agents are seeking improved token efficiency, reduced latency, and dependable performance. To address these needs, we are launching our Flash series models designed for optimal efficiency and quality, enhancing agentic workflows. Building on the capabilities of Gemini 3.5 Flash, we are excited to introduce the new Gemini models:
Gemini 3.6 Flash: Serving as our primary model, it enhances coding, knowledge work, and multimodal capabilities. The Artificial Analysis Index indicates that it achieves a 17% reduction in output token usage compared to the 3.5 Flash version. In specific benchmarks, such as DeepSWE by Datacurve, improvements can reach as high as 65%, all while maintaining a lower cost per output token.
Gemini 3.5 Flash-Lite: As the fastest and most economical model within the 3.5 range, this version can produce 350 output tokens per second, according to the Artificial Analysis Index. It also shows significant improvements over earlier Flash-Lite models in agentic workflows.
Gemini 3.5 Flash Cyber in CodeMender: To support effective cybersecurity solutions, a careful combination of a specialized, highly efficient model with an agent infrastructure is essential. We’re introducing a new cyber-focused model alongside our CodeMender code security agent, which provides competitive performance at the cutting edge.
Looking ahead, we are testing Gemini 3.5 Pro with select partners and anticipate a wider release as soon as it's ready. Concurrently, our team is ramping up for the next evolution of our models. We’ve initiated our most extensive pre-training run to date for Gemini 4 and are eager about the results we are seeing.
Gemini 3.6 Flash is designed with the feedback from developers and customers of 3.5 Flash in mind. This model not only enhances performance in coding and knowledge work, but it also significantly boosts token efficiency. The Artificial Analysis Index shows that 3.6 Flash uses 17% fewer output tokens than its predecessor. Moreover, it simplifies multi-step workflows by requiring fewer reasoning steps and tool calls.
This upgraded efficiency comes with a lower price point compared to 3.5 Flash. At a cost of $1.50 for every 1 million input tokens and $7.50 for 1 million output tokens, Gemini 3.6 Flash reduces the overall expense for agentic tasks, making it more affordable to develop and operate AI agents.



