At the recent I/O event in May, Google showcased an important update to its AI models with the introduction of Gemini 3.5 Flash, and now the company has unveiled three additional AI versions, including its inaugural Olympus model designed specifically for cybersecurity. Notably, these new offerings do not include the long-awaited Gemini 3.5 Pro, originally scheduled for a June debut.
Gemini 3.5 Flash, which initially received much attention, has already been retired in favor of Gemini 3.6 Flash. Google claims that this latest iteration is slightly more advanced, enhances coding capabilities, and boasts improved multimodal features.
The adjustments to 3.6 Flash were reportedly made in response to user insights regarding the 3.5 model, which fell short of Google’s expectations for code generation performance. This could reflect the company's increased emphasis on efficiency as organizations grow increasingly concerned about AI-related costs.
In coding evaluation, known as the DeepSWE test, Gemini 3.6 Flash improves its performance to 49 percent compared to the 37 percent achieved by its predecessor, 3.5 Flash. The updated model also integrates computer use as a fundamental component of the Gemini API, resulting in a modest increase in performance, elevating the OSWorld test score to 83 percent from 78.4 percent in 3.5.
Google has intensified its focus on efficiency with Gemini 3.6 Flash, and despite only slight improvements in benchmarks, this model consumes approximately 17 percent fewer tokens. It promises to execute agentic workflows more accurately, with fewer steps and reduced token usage, potentially resulting in significant savings for developers and the company alike. The new API pricing is set at $1.50 for every million input tokens and $7.50 for output tokens, a notable decrease from the previous $1.50 and $9 for 3.5 Flash.
In addition to these developments, Google has expanded the 3.5 series with the introduction of Gemini 3.5 Flash Lite and 3.5 Flash Cyber. The Flash Lite version emerges as the company's most efficient AI, achieving an impressive rate of 350 tokens per second. Google positions this model as cost-effective for scaling agentic systems. While the pricing is set at $0.30 per million input tokens and $2.50 for output tokens—slightly higher than the earlier 3.1 Flash Lite model, which was priced at $0.25 and $1.50—it still offers a competitive option relative to last year's frontier models.



