On July 21, 2026, Google released three new Gemini models at the cheaper, faster end of its lineup while its flagship Gemini 3.5 Pro remained in limited testing, missing the June target announced at I/O. The company also revealed it has begun “the most ambitious pre-training run yet” for Gemini 4.
The new models—Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber—are aimed at developers building production AI agents who need higher token efficiency, lower latency, and more reliable performance. The timing signals Google is prioritizing volume and speed over top-end reasoning, even as competitors push forward with frontier models.
What Happened
Gemini 3.6 Flash is the direct successor to 3.5 Flash. Google says it uses 17% fewer output tokens on the Artificial Analysis Index and takes fewer reasoning steps and tool calls to complete multi-step tasks. It is priced at $1.50 per million input tokens and $7.50 per million output tokens, cheaper than the previous Flash’s $9 per million output. Its knowledge cutoff moves from January 2025 to March 2026.
Gemini 3.5 Flash-Lite is described as the fastest, most cost-effective 3.5-class model, delivering 350 output tokens per second. It significantly outperforms prior Flash-Lite generations in agentic workflows.
Gemini 3.5 Flash Cyber is a security-tuned variant built for cybersecurity applications, developed under the name CodeMender.
Meanwhile, Gemini 3.5 Pro—which Google promised at I/O in May would arrive in June—remains in limited testing with partners. The company now says it “will ship when it’s ready.” This delay, combined with the three Flash launches, suggests Google is leaning hard into the efficient, high-volume tier while its flagship work continues behind closed doors. In the same announcement, Google confirmed it has already started pre-training for Gemini 4.
My Take
Google is playing a smart game. By shipping three Flash models simultaneously—including a niche Cyber variant—they’re showing they can iterate quickly on the efficiency side while the core model (3.5 Pro) still cooks. The 17% token reduction on 3.6 Flash is meaningful for developers running production agents at scale; cost savings compound fast. The Flash-Lite’s 350 tokens/sec is a clear message for real-time or streaming use cases.
But the delay of 3.5 Pro is worrying. It implies either the model isn’t hitting the necessary quality bar, or Google is rethinking its architecture mid-stream. Either way, competitors like Anthropic, OpenAI, and Meta are not standing still. Google’s mention of Gemini 4 pre-training is a hedge—it tells the market “we’re thinking long-term,” but developers need a capable flagship now. For now, 3.6 Flash is a solid workhorse, but it’s not the top-tier model the community was waiting for.
What to Watch
- Pricing pressure: At $7.50 per million output tokens, Google is undercutting its own previous Flash pricing. Expect competitors to respond with cost cuts.
- 3.5 Pro timeline: If Pro doesn’t ship within Q3 2026, Google risks losing enterprise customers who need state-of-the-art reasoning.
- Gemini 4 pre-training: The scale of “most ambitious run yet” hints at a major architectural shift—possibly a mixture-of-experts or retrieval-augmented approach baked in from the start.
