Google Announces Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber: Everything You Need to Know
Google just launched Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber. Pricing, benchmarks, availability, safety and the Gemini 4 roadmap — full breakdown.
Three new Flash models in one day
On July 21, 2026, Google announced three new Gemini models in a single drop: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. The through-line is unmistakable — every one of them targets the same problem developers keep raising: production AI agents need higher token efficiency, lower latency, and more reliable performance, not just bigger headline scores.
This post breaks down what each model is for, what it costs, where you can use it today, and what Google teased about the road ahead (spoiler: Gemini 3.5 Pro is in partner testing, and the Gemini 4 pre-training run has already started). All figures come from the official announcement.
The lineup at a glance
| Model | Role | Headline stat | Access |
|---|---|---|---|
| Gemini 3.6 Flash | Workhorse: coding, knowledge work, multimodal | 17% fewer output tokens than 3.5 Flash, at a lower price | Generally available today |
| Gemini 3.5 Flash-Lite | Fastest, most cost-effective 3.5-class model | 350 output tokens/s (Artificial Analysis) | Generally available today |
| Gemini 3.5 Flash Cyber | Security specialist for finding & fixing vulnerabilities | Frontier-competitive on CyberGym inside CodeMender | Limited-access pilot only |
Gemini 3.6 Flash: the new workhorse
3.6 Flash builds directly on developer feedback from 3.5 Flash. The pitch is unusual for a model launch: it’s not just better, it’s more frugal. On the Artificial Analysis Index it consumes 17% fewer output tokens than 3.5 Flash, and on some benchmarks — like DeepSWE by Datacurve — Google observed reductions of up to 65%. It also needs fewer reasoning steps and tool calls to finish multi-step workflows, which compounds into real savings on agentic tasks.

Benchmark highlights versus 3.5 Flash:
| Benchmark | Gemini 3.6 Flash | Gemini 3.5 Flash |
|---|---|---|
| DeepSWE (agentic coding) | 49% | 37% |
| MLE Bench (ML research) | 63.9% | 49.7% |
| OSWorld-Verified (computer use) | 83.0% | 78.4% |
| GDPval-AA v2 (knowledge work) | 1421 | 1349 |
Two more launch-day details matter for builders:
- Computer use is now a built-in client-side tool via the Gemini API and Gemini Enterprise — no separate scaffolding required.
- Early customers include Figma, Harvey, Hebbia, and JetBrains; Hebbia and Harvey specifically call out multimodal strengths like document parsing, chart and data analysis, and report drafting.
Google’s demo of side-by-side token efficiency on an OSWorld task is worth watching: 3.6 Flash vs 3.5 Flash token-efficiency demo (video).
For the full head-to-head, see our Gemini 3.6 Flash vs 3.5 Flash deep dive.
Gemini 3.5 Flash-Lite: built to scale agentic workflows
Flash-Lite is the volume play. It’s the fastest model in the 3.5 series at 350 output tokens per second (Artificial Analysis), aimed at agentic search, document processing, and any pipeline where throughput is the bottleneck.
The quality jump over 3.1 Flash-Lite is dramatic — Terminal-Bench 2.1 at 54% vs 31%, GDM-MRCR v2 long-context at 72.2% vs 60.1%, and GDPval-AA v2 at 1140 vs 642. Remarkably, it even beats the bigger Gemini 3 Flash on several agentic evals, including SWE-Bench Pro (54.2% vs 49.6%) and OSWorld-Verified (74.0% vs 65.1%).
It ships with configurable thinking levels — run minimal/low for latency-critical high-volume tasks, or raise the level for multi-step subagent workloads — and computer use is built in here too. Early customers include Ashler, Palo Alto Networks, and Ramp, and the model is rolling out inside Google Search and the Gemini app.
Our 3.5 Flash-Lite agentic workflows guide covers thinking levels, cost math, and model selection in depth.
Pricing: cheaper at both tiers
| Model | Input ($/1M tokens) | Output ($/1M tokens) | Notes |
|---|---|---|---|
| Gemini 3.6 Flash | $1.50 | $7.50 | Lower than 3.5 Flash, plus ~17% fewer output tokens |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | Best price-to-performance for high-throughput traffic |
| Gemini 3.5 Flash Cyber | — | — | Not publicly priced; limited-access pilot via CodeMender |
The real cost story is the multiplication: 3.6 Flash charges less per token and emits fewer tokens per task, so the effective cost per agentic task drops more than the price sheet suggests.
Gemini 3.5 Flash Cyber: a specialist behind a velvet rope
The third model is different. 3.5 Flash Cyber is built on 3.5 Flash and fine-tuned for finding and fixing security vulnerabilities at a lower price per token than larger models. It runs inside CodeMender, Google’s code-security agent, where multiple 3.5 Flash Cyber agents work together to produce a single combined report — reaching competitive frontier performance on the CyberGym benchmark.

Because the technology is dual-use, Google is deliberately not releasing it broadly: it will be exclusively available to governments and trusted partners via CodeMender as part of a limited-access pilot. The goal is to give frontline defenders a head start on patching critical vulnerabilities before they’re exploited.
Safety: fewer jailbreaks, fewer refusals
3.6 Flash ships with enhanced Frontier Safety safeguards covering CBRN (Chemical, Biological, Radiological, Nuclear) and cyber-offense misuse. Google says the model is substantially more resistant to jailbreaks while simultaneously trained to minimize refusals for benign uses — the combination developers actually want, rather than a blunt refusal-everything posture.
Where you can use them today
| Surface | Gemini 3.6 Flash | Gemini 3.5 Flash-Lite |
|---|---|---|
| Gemini API (Google AI Studio, Android Studio) | ✅ | ✅ |
| Google Antigravity | ✅ | — |
| Gemini Enterprise Agent Platform | ✅ | ✅ |
| Gemini Enterprise app | ✅ | — |
| Gemini app | ✅ | ✅ |
| Google Search | — | ✅ Rolling out |
| CodeMender (3.5 Flash Cyber) | Limited pilot: governments & trusted partners only |
The roadmap teaser: 3.5 Pro and Gemini 4
Two forward-looking notes buried in the announcement deserve attention:
- Gemini 3.5 Pro is currently testing with partners, with broad availability planned “as soon as it’s ready.”
- Google has started its most ambitious pre-training run yet — for Gemini 4 — and says it’s “excited by the progress.”
Read together, the Flash-tier releases look like the efficiency layer being locked in before the next frontier push.
Bottom line
If you build agents, this is a rare launch where the right move is obvious: 3.6 Flash for quality-sensitive agentic work, 3.5 Flash-Lite for high-volume throughput, both cheaper than what they replace. Flash Cyber, meanwhile, signals where Google thinks specialized fine-tunes are headed — powerful, narrow, and deliberately gated.