AI

Google Releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

Google's new Gemini 3.6 Flash cuts output token usage by 17% while improving coding and agentic performance, alongside a faster 3.5 Flash-Lite and a security-focused 3.5 Flash Cyber model.

Google released three new Gemini models: 3.6 Flash, which cuts output tokens by 17% while improving coding and agentic performance at lower cost; 3.5 Flash-Lite, delivering 350 tokens per second for high-throughput workloads; and 3.5 Flash Cyber, a security-focused model restricted to government and trusted partners. Gemini 3.5 Pro remains in testing, and Google has begun pre-training for Gemini 4.

Three New Models, One Focus: Agentic Efficiency

Google has released three new models in its Gemini Flash series, each targeting a specific bottleneck in production AI agent deployment. The releases are Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber, alongside confirmation that Gemini 3.5 Pro remains in partner testing and that work on Gemini 4 has begun.

Gemini 3.6 Flash: The Workhorse Upgrade

Gemini 3.6 Flash is positioned as the primary model for developers building agents at scale. It delivers improvements in coding, knowledge work, and multimodal performance while reducing token consumption. According to the Artificial Analysis Index, it uses 17% fewer output tokens than 3.5 Flash, with some benchmarks like DeepSWE showing reductions up to 65%. Despite the efficiency gains, it outperforms its predecessor across key metrics:

  • DeepSWE: 49% vs. 37% (fewer unwanted edits, reduced execution loops)
  • MLE Bench: 63.9% vs. 49.7% (ML research tasks)
  • OSWorld-Verified: 83.0% vs. 78.4% (computer use capabilities)
  • GDPval-AA v2: 1421 vs. 1349 (knowledge work)

The model is priced at 1.50permillioninputtokensand1.50 per million input tokens and 7.50 per million output tokens, lower than 3.5 Flash, making it cheaper to run despite the performance gains. Computer use is now a built-in client-side tool via the Gemini API and Gemini Enterprise.

Google also highlights enhanced safeguards in Chemical, Biological, Radiological, and Nuclear (CBRN) domains and cyber offense misuse, with the model trained to minimize refusals for beneficial uses while resisting jailbreaks.

3.5 Flash-Lite: Speed at Scale

For high-throughput applications, 3.5 Flash-Lite is the fastest model in the 3.5 series, delivering 350 output tokens per second according to Artificial Analysis. Priced at 0.30permillioninputtokensand0.30 per million input tokens and 2.50 per million output tokens, it offers a strong price-to-performance ratio for document processing and agentic search workloads.

The model significantly outperforms 3.1 Flash-Lite across benchmarks:

  • Terminal-Bench 2.1: 54% vs. 31%
  • GDM-MRCR v2: 72.2% vs. 60.1%
  • GDPval-AA v2: 1140 vs. 642

Notably, it even surpasses the larger 3 Flash on some agentic and coding evaluations, including SWE-Bench Pro (54.2% vs. 49.6%) and OSWorld-Verified (74.0% vs. 65.1%). It includes configurable thinking levels, allowing developers to prioritize low-latency execution or engage deeper reasoning for multi-step subagent tasks.

3.5 Flash Cyber: Security-Focused and Controlled

The most specialized release is 3.5 Flash Cyber, a fine-tuned variant built on 3.5 Flash and paired with Google's CodeMender code security agent. It is designed for finding and fixing cybersecurity vulnerabilities at scale, reaching competitive frontier performance on the CyberGym benchmark.

Given the dual-use nature of the technology, Google is taking a restrictive approach to deployment. The model will be exclusively available to governments and trusted partners via a limited-access pilot program, intended to give frontline defenders an advantage in vulnerability discovery while mitigating broader misuse risks.

The Road Ahead

Google confirmed that Gemini 3.5 Pro is currently in testing with partners and will be released broadly when ready. More significantly, the company announced it has started its "most ambitious pre-training run yet" for Gemini 4, signaling that the current releases are interim steps toward a larger generational leap.

The pricing and performance positioning of 3.6 Flash and 3.5 Flash-Lite suggests Google is aggressively targeting the developer market for agentic applications, where token efficiency and latency often matter more than raw capability. The cybersecurity model's restricted release, by contrast, reflects growing industry awareness that the most capable AI systems for security tasks require careful access controls. Together, the three releases cover the mainstream, high-volume, and specialized security segments of the agentic AI market.