Google shipped three new Gemini models. The one everyone's waiting for still isn't here.
Gemini 3.6 Flash, 3.5 Flash-Lite, and a government-only 3.5 Flash Cyber model landed July 21. Gemini 3.5 Pro remains delayed over internal benchmark shortfalls.
Google shipped 3 new Gemini models on July 21. The flagship 3.5 Pro is still stuck in testing.
Google shipped three new Gemini models on July 21: a cheaper, faster 3.6 Flash, a low latency 3.5 Flash-Lite, and a security focused 3.5 Flash Cyber restricted to governments and trusted partners. Gemini 3.5 Pro, the model most developers are actually waiting for, is still not out.
What actually changed
Gemini 3.6 Flash cuts output token cost from $9.00 to $7.50 per million tokens, a 17 percent reduction, and improves computer use accuracy from 78.4 to 83.0 percent on OSWorld-Verified. Gemini 3.5 Flash-Lite is tuned for raw throughput: 350 output tokens per second, and it jumped from 31 to 54 percent on Terminal-Bench 2.1. Figma, Harvey, JetBrains, and Ashby are named as early users in Google's own post.
The one that is not for you
Gemini 3.5 Flash Cyber is a version of 3.5 Flash specialized for finding and patching security vulnerabilities, and Google is gating access to governments and trusted partners rather than shipping it broadly. That is a deliberate choice given how dual use a vulnerability finding model is, the same logic behind Cisco gating its own Antares models to vetted downloads this week.
Why 3.5 Pro is still missing
Per Bloomberg reporting cited by TechCrunch, Gemini 3.5 Pro "struggled to meet internal performance goals." Google DeepMind's Logan Kilpatrick says it is in partner testing now, and the company has already begun pretraining Gemini 4. A frontier lab shipping three secondary models while its flagship stays in the oven is a concrete data point on how much friction is actually involved in hitting a self imposed bar right now.
Why a build studio cares
The pricing and benchmark deltas here are the kind of numbers that actually move a model selection decision for an agentic build: cheaper output tokens and faster throughput matter more than another point of benchmark score for most client work. The Pro delay is the more interesting signal though. If Google's own bar is hard to clear, that is worth knowing before planning a build around whatever ships next.
Next step: read Google's own announcement and TechCrunch's coverage of the Pro delay. If you're picking a model for an agentic workload and want the current tradeoffs, write to us at hello@gattyworks.com.