Gemini Omni
Back to all articles
8 min read

Google Announces Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber: Everything You Need to Know

Google just launched Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber. Pricing, benchmarks, availability, safety and the Gemini 4 roadmap — full breakdown.

Gemini 3.6 FlashGemini 3.5 Flash-LiteGemini 3.5 Flash CyberModelsAnnouncement2026

Three new Flash models in one day

On July 21, 2026, Google announced three new Gemini models in a single drop: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. The through-line is unmistakable — every one of them targets the same problem developers keep raising: production AI agents need higher token efficiency, lower latency, and more reliable performance, not just bigger headline scores.

This post breaks down what each model is for, what it costs, where you can use it today, and what Google teased about the road ahead (spoiler: Gemini 3.5 Pro is in partner testing, and the Gemini 4 pre-training run has already started). All figures come from the official announcement.

The lineup at a glance

ModelRoleHeadline statAccess
Gemini 3.6 FlashWorkhorse: coding, knowledge work, multimodal17% fewer output tokens than 3.5 Flash, at a lower priceGenerally available today
Gemini 3.5 Flash-LiteFastest, most cost-effective 3.5-class model350 output tokens/s (Artificial Analysis)Generally available today
Gemini 3.5 Flash CyberSecurity specialist for finding & fixing vulnerabilitiesFrontier-competitive on CyberGym inside CodeMenderLimited-access pilot only

Gemini 3.6 Flash: the new workhorse

3.6 Flash builds directly on developer feedback from 3.5 Flash. The pitch is unusual for a model launch: it’s not just better, it’s more frugal. On the Artificial Analysis Index it consumes 17% fewer output tokens than 3.5 Flash, and on some benchmarks — like DeepSWE by Datacurve — Google observed reductions of up to 65%. It also needs fewer reasoning steps and tool calls to finish multi-step workflows, which compounds into real savings on agentic tasks.

Gemini 3.6 Flash benchmark gains over 3.5 Flash across DeepSWE, MLE Bench, OSWorld-Verified and GDPval-AA v2

Benchmark highlights versus 3.5 Flash:

BenchmarkGemini 3.6 FlashGemini 3.5 Flash
DeepSWE (agentic coding)49%37%
MLE Bench (ML research)63.9%49.7%
OSWorld-Verified (computer use)83.0%78.4%
GDPval-AA v2 (knowledge work)14211349

Two more launch-day details matter for builders:

  • Computer use is now a built-in client-side tool via the Gemini API and Gemini Enterprise — no separate scaffolding required.
  • Early customers include Figma, Harvey, Hebbia, and JetBrains; Hebbia and Harvey specifically call out multimodal strengths like document parsing, chart and data analysis, and report drafting.

Google’s demo of side-by-side token efficiency on an OSWorld task is worth watching: 3.6 Flash vs 3.5 Flash token-efficiency demo (video).

For the full head-to-head, see our Gemini 3.6 Flash vs 3.5 Flash deep dive.

Gemini 3.5 Flash-Lite: built to scale agentic workflows

Flash-Lite is the volume play. It’s the fastest model in the 3.5 series at 350 output tokens per second (Artificial Analysis), aimed at agentic search, document processing, and any pipeline where throughput is the bottleneck.

The quality jump over 3.1 Flash-Lite is dramatic — Terminal-Bench 2.1 at 54% vs 31%, GDM-MRCR v2 long-context at 72.2% vs 60.1%, and GDPval-AA v2 at 1140 vs 642. Remarkably, it even beats the bigger Gemini 3 Flash on several agentic evals, including SWE-Bench Pro (54.2% vs 49.6%) and OSWorld-Verified (74.0% vs 65.1%).

It ships with configurable thinking levels — run minimal/low for latency-critical high-volume tasks, or raise the level for multi-step subagent workloads — and computer use is built in here too. Early customers include Ashler, Palo Alto Networks, and Ramp, and the model is rolling out inside Google Search and the Gemini app.

Our 3.5 Flash-Lite agentic workflows guide covers thinking levels, cost math, and model selection in depth.

Pricing: cheaper at both tiers

ModelInput ($/1M tokens)Output ($/1M tokens)Notes
Gemini 3.6 Flash$1.50$7.50Lower than 3.5 Flash, plus ~17% fewer output tokens
Gemini 3.5 Flash-Lite$0.30$2.50Best price-to-performance for high-throughput traffic
Gemini 3.5 Flash CyberNot publicly priced; limited-access pilot via CodeMender

The real cost story is the multiplication: 3.6 Flash charges less per token and emits fewer tokens per task, so the effective cost per agentic task drops more than the price sheet suggests.

Gemini 3.5 Flash Cyber: a specialist behind a velvet rope

The third model is different. 3.5 Flash Cyber is built on 3.5 Flash and fine-tuned for finding and fixing security vulnerabilities at a lower price per token than larger models. It runs inside CodeMender, Google’s code-security agent, where multiple 3.5 Flash Cyber agents work together to produce a single combined report — reaching competitive frontier performance on the CyberGym benchmark.

Gemini 3.5 Flash Cyber CyberGym benchmark results inside CodeMender

Because the technology is dual-use, Google is deliberately not releasing it broadly: it will be exclusively available to governments and trusted partners via CodeMender as part of a limited-access pilot. The goal is to give frontline defenders a head start on patching critical vulnerabilities before they’re exploited.

Safety: fewer jailbreaks, fewer refusals

3.6 Flash ships with enhanced Frontier Safety safeguards covering CBRN (Chemical, Biological, Radiological, Nuclear) and cyber-offense misuse. Google says the model is substantially more resistant to jailbreaks while simultaneously trained to minimize refusals for benign uses — the combination developers actually want, rather than a blunt refusal-everything posture.

Where you can use them today

SurfaceGemini 3.6 FlashGemini 3.5 Flash-Lite
Gemini API (Google AI Studio, Android Studio)
Google Antigravity
Gemini Enterprise Agent Platform
Gemini Enterprise app
Gemini app
Google Search✅ Rolling out
CodeMender (3.5 Flash Cyber)Limited pilot: governments & trusted partners only

The roadmap teaser: 3.5 Pro and Gemini 4

Two forward-looking notes buried in the announcement deserve attention:

  1. Gemini 3.5 Pro is currently testing with partners, with broad availability planned “as soon as it’s ready.”
  2. Google has started its most ambitious pre-training run yet — for Gemini 4 — and says it’s “excited by the progress.”

Read together, the Flash-tier releases look like the efficiency layer being locked in before the next frontier push.

Bottom line

If you build agents, this is a rare launch where the right move is obvious: 3.6 Flash for quality-sensitive agentic work, 3.5 Flash-Lite for high-volume throughput, both cheaper than what they replace. Flash Cyber, meanwhile, signals where Google thinks specialized fine-tunes are headed — powerful, narrow, and deliberately gated.