Google has unveiled three new additions to its Gemini AI lineup: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. While the launch introduces faster, more efficient, and more specialized AI models, one highly anticipated release is still missing Gemini 3.5 Pro.
The latest updates focus on improving coding performance, lowering AI operating costs, and expanding Google’s AI capabilities into cybersecurity. According to Google’s announcement, Gemini 3.6 Flash delivers stronger coding performance while reducing output token usage by 17% compared to Gemini 3.5 Flash. Meanwhile, Flash-Lite prioritizes speed and affordability, and Flash Cyber introduces Google’s first AI model specifically designed for cybersecurity workflows.
Although these releases represent a significant step forward for Google’s AI ecosystem, developers are still waiting for Gemini 3.5 Pro, which remains in partner testing without a confirmed public release date.
Gemini 3.6 Flash builds directly on developer and customer feedback from 3.5 Flash. 3.6 Flash not only delivers a step up in coding and knowledge work, but it does this while meaningfully improving token efficiency.
Gemini 3.6 Flash: Better Performance With Lower Costs
Gemini 3.6 Flash replaces Gemini 3.5 Flash as Google’s default high-performance Flash model. It focuses on improving coding quality, knowledge-based tasks, and overall efficiency while keeping operating costs under control. According to Google, the model was refined using feedback from developers who wanted better code generation and more reliable performance in production environments.
Key improvements include:
- Better coding capabilities
- Stronger knowledge-work performance
- Improved multimodal understanding
- Lower output token usage
- Reduced API costs
- Enhanced safety protections for cyber-related and CBRN misuse
One of the biggest improvements is token efficiency. Google says Gemini 3.6 Flash consumes approximately 17% fewer output tokens than Gemini 3.5 Flash while producing comparable results. On the DeepSWE coding benchmark, that efficiency reportedly increases to around 65%, allowing developers to complete complex coding tasks with significantly fewer generated tokens. For organizations running AI-powered applications at scale, lower token usage translates directly into lower infrastructure costs.
Gemini 3.5 Flash-Lite Prioritizes Speed
While Gemini 3.6 Flash balances performance and efficiency, Gemini 3.5 Flash-Lite is designed for speed and high-volume workloads. Google says Flash-Lite can generate approximately 350 output tokens per second, making it the fastest model in the current Flash family. Rather than maximizing reasoning ability, it focuses on delivering responses quickly and at a lower cost.
The model is intended for applications such as:
- AI-powered search experiences
- Customer support assistants
- Large-scale AI agents
- Content classification
- High-volume enterprise automation
Google is also deploying Flash-Lite across several of its own services, including Google Search, where it powers AI-generated experiences that require rapid responses at massive scale. The company views Flash-Lite as the ideal balance between speed, quality, and affordability for businesses handling millions of AI requests each day.
Gemini 3.5 Flash Cyber Brings AI to Cybersecurity
The third addition to Google’s lineup is Gemini 3.5 Flash Cyber, the company’s first language model built specifically for cybersecurity. Unlike standard conversational AI models, Flash Cyber is designed to identify software vulnerabilities and support security investigations through Google’s CodeMender platform.
Instead of analyzing an entire codebase in a single pass, CodeMender distributes different sections of the code across multiple Flash Cyber instances running simultaneously. The results are then combined into a unified vulnerability report, enabling faster and more comprehensive security analysis.
Google reports that, during testing on Chrome’s V8 engine, Flash Cyber identified:
- 55 confirmed unique vulnerabilities
- Compared with 47 found by standard Gemini 3.5 Flash
- And 36 identified by Claude Opus 4.6
The model reportedly uncovered several vulnerabilities that competing models failed to detect. Because of its advanced security capabilities, Flash Cyber will not be publicly available. Google plans to release it initially through a limited pilot program restricted to governments and trusted enterprise partners.
Gemini 3.5 Pro Remains Delayed
Despite launching three new AI models, Google has yet to release Gemini 3.5 Pro. The company originally suggested that the flagship model would arrive shortly after Google I/O in May. However, months later, it remains in partner testing with no official public release date.
Reports suggest the delay is linked to Google’s internal quality standards, with additional testing continuing before a broader rollout. According to Google, Gemini 3.5 Pro will be released once it meets the company’s performance expectations rather than following a fixed launch schedule.
Why Token Efficiency Matters
One of the most significant improvements in Gemini 3.6 Flash is its reduced token consumption. For AI developers, tokens directly influence operating costs. Every request sent to an AI model and every generated response consumes tokens that are billed through API pricing. A reduction of 17% in output tokens can create substantial savings for businesses processing millions of requests each month.
Lower token usage also offers additional operational benefits, including:
- Lower infrastructure expenses
- Faster AI responses
- More efficient agent workflows
- Reduced reasoning overhead
- Lower costs for enterprise AI deployments
If these efficiency improvements continue to perform well outside benchmark environments, they could become one of Gemini 3.6 Flash’s strongest competitive advantages.

Google’s Three-Model Strategy Signals a Shift
The launch of three specialized Flash models reflects a broader change in Google’s AI strategy. Rather than relying on a single model for every scenario, Google is separating its AI offerings based on different priorities:
| Model | Primary Focus |
|---|---|
| Gemini 3.6 Flash | Coding, knowledge work, and efficient general AI |
| Gemini 3.5 Flash-Lite | High-speed, low-cost AI at scale |
| Gemini 3.5 Flash Cyber | Cybersecurity and vulnerability detection |
This approach allows developers to choose a model based on workload rather than paying for capabilities they may not need.
Competition in the AI Race Continues
Google’s latest releases arrive during an increasingly competitive period for the AI industry. Major AI companies continue introducing new flagship models at a rapid pace, making performance, efficiency, and cost optimization key areas of competition.
While many competitors emphasize reasoning ability and benchmark scores, Google’s latest announcements suggest that infrastructure efficiency and specialized AI models are becoming equally important. Whether Gemini 3.6 Flash delivers the same gains in real-world deployments as it does in benchmark testing remains a question developers will answer through production workloads.
Key Takeaways
Google’s latest Gemini update is less about introducing a single breakthrough model and more about offering specialized AI solutions for different needs.
The announcement introduces:
- Gemini 3.6 Flash for stronger coding, knowledge work, and improved token efficiency.
- Gemini 3.5 Flash-Lite for high-speed, cost-effective AI workloads.
- Gemini 3.5 Flash Cyber for advanced cybersecurity and vulnerability detection.
At the same time, the continued absence of Gemini 3.5 Pro indicates that Google is prioritizing product readiness over rushing its flagship model to market.
For developers building AI agents, enterprise applications, or security tools, these releases provide new options that emphasize efficiency, scalability, and specialized performance rather than raw benchmark leadership alone.
Your Queries
How much does Gemini 3.6 Flash reduce output token usage?
According to Google, Gemini 3.6 Flash uses approximately 17% fewer output tokens than Gemini 3.5 Flash for comparable workloads.
What are the three new Gemini models?
Google introduced Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber, each designed for different AI workloads.
Why is token efficiency important?
Lower token usage reduces API costs, improves scalability, and helps organizations operate AI systems more efficiently.
What improvements does Gemini 3.6 Flash offer?
The model delivers stronger coding performance, improved knowledge work, lower token consumption, enhanced safety guardrails, and better cost efficiency.
Is Gemini 3.5 Flash Cyber publicly available?
No. Google is initially making Flash Cyber available only to governments and trusted partners through a limited pilot program.
When will Gemini 3.5 Pro be released?
Google has not announced an official launch date. The model remains in partner testing.
Where can developers access Gemini 3.6 Flash and Flash-Lite?
Both models are available through Google AI Studio, the Gemini API, and Gemini Enterprise, while Flash-Lite is also being integrated into Google Search experiences.