Technology

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

Google launched three new Gemini models built for production AI agents, with 3.6 Flash cutting token costs by 17% and a specialized cybersecurity variant for defenders.

Google released three new Gemini models: 3.6 Flash, which improves token efficiency by 17% over its predecessor while scoring higher on coding and computer-use benchmarks; 3.5 Flash-Lite, a high-speed budget option at 350 tokens per second; and 3.5 Flash Cyber, a restricted-access model fine-tuned for cybersecurity vulnerability detection through the CodeMender agent.

Google has expanded its Gemini Flash family with three new models, each targeting a different slice of the growing market for production AI agents. The releases include Gemini 3.6 Flash, a workhorse model for coding and knowledge tasks; 3.5 Flash-Lite, optimized for speed and throughput; and 3.5 Flash Cyber, a specialized variant fine-tuned for cybersecurity operations.

Gemini 3.6 Flash is the headline release. Google says it consumes 17% fewer output tokens than its 3.5 Flash predecessor while delivering better results on coding, multimodal reasoning, and computer-use benchmarks. On DeepSWE, a software engineering evaluation, 3.6 Flash hit 49% versus 3.5 Flash's 37%. On OSWorld-Verified, a computer-use test, it scored 83.0% against 78.4%. The model is priced at 1.50permillioninputtokens∗∗and∗∗1.50 per million input tokens** and **7.50 per million output tokens, a reduction from 3.5 Flash's rates.

Google also claims 3.6 Flash takes fewer reasoning steps and tool calls to complete multi-step workflows, a meaningful improvement for agentic systems where latency compounds across chained operations. The model ships with enhanced safeguards against jailbreaks in CBRN and cyber offense domains, while Google says it has been trained to minimize unnecessary refusals for legitimate use cases.

Gemini 3.5 Flash-Lite is positioned as the fastest and most cost-effective model in the 3.5 series. It outputs at 350 tokens per second according to the Artificial Analysis Index, priced at 0.30permillioninputtokens∗∗and∗∗0.30 per million input tokens** and **2.50 per million output tokens. Despite the low cost, Google says it outperforms 3.1 Flash-Lite by wide margins on Terminal-Bench 2.1 (54% versus 31%) and long-context tasks (72.2% versus 60.1%). It even beats the full 3 Flash on some agentic and coding evaluations.

The most unusual release is Gemini 3.5 Flash Cyber, a model fine-tuned specifically for finding, validating, and patching software vulnerabilities. Built on top of 3.5 Flash and integrated into Google's CodeMender security agent, it is designed to be deployed in multiple parallel agents that scan codebases and produce combined vulnerability reports. Google says it reached competitive performance on the CyberGym benchmark against significantly larger models.

Because of the dual-use nature of offensive cybersecurity capabilities, 3.5 Flash Cyber will be available only to governments and trusted partners through a limited-access pilot program. Google has not announced a timeline for broader availability.

All three models are available to developers through the Gemini API and Google AI Studio, with 3.6 Flash also accessible in the Gemini app. Enterprise users can access them through Vertex AI. Google also confirmed that Gemini 3.5 Pro is currently in partner testing, and that the company has begun what it calls its "most ambitious pre-training run yet" for Gemini 4.

The release pattern is telling. Instead of a single flagship model, Google is segmenting its lineup by workload: one for quality and efficiency, one for raw speed and cost, and one for a specialized vertical where safety considerations demand restricted access. This reflects a maturing market where developers are no longer asking "which model is best?" but "which model is best for this specific task at this specific cost?"